Shijie Peng (彭士杰)
PhD Student in Computer Science, UCAS & SIAT, CAS
UCAS & SIAT, CAS
I am a PhD student in Computer Science at the University of Chinese Academy of Sciences (UCAS), advised by Prof. Kejiang Ye (叶可江), and affiliated with the Shenzhen Institutes of Advanced Technology (SIAT), Chinese Academy of Sciences.
My research focuses on systems for AI, with an emphasis on making large-model inference efficient, elastic, and reliable in real-world cloud environments. I study the interaction between model execution, serving runtimes, scheduling, and heterogeneous hardware.
Research Interests
- LLM serving systems: architecture and runtime design for low-latency, high-throughput inference
- Serverless and cloud systems: elastic scheduling and resource management for AI-native workloads
- Operating and distributed systems: performance optimization across heterogeneous clusters
- Hardware-software co-design for AI: practical acceleration strategies for production inference pipelines
Research & Collaborations
I collaborate with academic and industry teams on large-scale AI infrastructure, especially model serving, scheduling, and end-to-end system efficiency. My work has appeared in venues including EuroSys, ICPP, TPDS, SoCC, ICDCS, and ICWS; see the publications page for the complete list.
News
| Aug 24, 2026 | Our paper “Valve” has been conditionally accepted to EuroSys’27! |
|---|---|
| Jul 06, 2026 | Our paper “ReliefServe” has been accepted to ICPP’26! |
| Sep 10, 2025 | Our paper “FlexPipe” has been accepted to EuroSys’26! |
Selected Publications
- EuroSys27Valve: Risk-Bounded Offline Execution in Online Generative ServingIn Proceedings of the 22nd European Conference on Computer Systems, 2027
- TPDSWorkload-Adapted Resource Allocation for LLM Distributed Serving in Serverless ClustersIEEE Transactions on Parallel and Distributed Systems, 2026
- ICWSPlanck: Optimizing LLM Inference Performance in Pipeline Parallelism with Fine-Grained SLO ConstraintIn Proceedings of the 31st IEEE International Conference on Web Services, 2024
- CSCWDEINS: Edge-Cloud Deep Model Inference with Network-Efficiency Schedule in ServerlessIn Proceedings of the 27th International Conference on Computer Supported Cooperative Work in Design, 2024