Shijie Peng (彭士杰)

PhD Student in Computer Science, UCAS & SIAT, CAS

37086559293.png

UCAS & SIAT, CAS

myzhibei@qq.com

I am a PhD student in Computer Science at the University of Chinese Academy of Sciences (UCAS), advised by Prof. Kejiang Ye (叶可江), and affiliated with the Shenzhen Institutes of Advanced Technology (SIAT), Chinese Academy of Sciences.

My research focuses on systems for AI, with an emphasis on making large-model inference efficient, elastic, and reliable in real-world cloud environments. I study the interaction between model execution, serving runtimes, scheduling, and heterogeneous hardware.

Research Interests

  • LLM serving systems: architecture and runtime design for low-latency, high-throughput inference
  • Serverless and cloud systems: elastic scheduling and resource management for AI-native workloads
  • Operating and distributed systems: performance optimization across heterogeneous clusters
  • Hardware-software co-design for AI: practical acceleration strategies for production inference pipelines

Research & Collaborations

I collaborate with academic and industry teams on large-scale AI infrastructure, especially model serving, scheduling, and end-to-end system efficiency. My work has appeared in venues including EuroSys, ICPP, TPDS, SoCC, ICDCS, and ICWS; see the publications page for the complete list.

News

Aug 24, 2026 Our paper “Valve” has been conditionally accepted to EuroSys’27!
Jul 06, 2026 Our paper “ReliefServe” has been accepted to ICPP’26!
Sep 10, 2025 Our paper “FlexPipe” has been accepted to EuroSys’26!

View all news →

Selected Publications

  1. EuroSys27
    Valve: Risk-Bounded Offline Execution in Online Generative Serving
    Yanying Lin , Haiying Shen , Shijie Peng, and 6 more authors
    In Proceedings of the 22nd European Conference on Computer Systems, 2027
  1. ICPP
    ReliefServe: Relieving GPU Pressure in Multi-Model Serving via Selective CPU Escape
    Shijie Peng , Yanying Lin , Chengzhi Lu, and 3 more authors
    In Proceedings of the 55th IEEE International Conference on Parallel Processing, 2026
  1. EuroSys
    FlexPipe: Adapting Dynamic LLM Serving Through Inflight Pipeline Refactoring in Fragmented Serverless Clusters
    Yanying Lin , Shijie Peng , Chengzhi Lu, and 2 more authors
    In Proceedings of the 21st European Conference on Computer Systems, 2026
  1. TPDS
    Workload-Adapted Resource Allocation for LLM Distributed Serving in Serverless Clusters
    Yanying Lin , Shijie Peng , Yanbo Li, and 4 more authors
    IEEE Transactions on Parallel and Distributed Systems, 2026
  1. Cluster
    ROCK: Serving Multimodal Models in Cloud with Heterogeneous-Aware Resource Orchestration for Thousands of LoRA Adapters
    Shuaipeng Wu , Yanying Lin , Shijie Peng, and 6 more authors
    In Proceedings of the 2025 IEEE International Conference on Cluster Computing, 2025
  1. IEEE TSC
    Serving LLM in Distributed GPU Cluster With Fine-Grain Pipeline Constraints
    Yanying Lin , Shijie Peng , Shuaipeng Wu, and 4 more authors
    IEEE Transactions on Services Computing, 2025
  1. ICDCS
    QUART: Latency-Aware FaaS System for Pipelining Large Model Inference
    Yanying Lin , Yanbo Li , Shijie Peng, and 5 more authors
    In Proceedings of the 44th IEEE International Conference on Distributed Computing Systems, 2024
  1. ICWS
    Planck: Optimizing LLM Inference Performance in Pipeline Parallelism with Fine-Grained SLO Constraint
    Yanying Lin , Shijie Peng , Shuaipeng Wu, and 4 more authors
    In Proceedings of the 31st IEEE International Conference on Web Services, 2024
  1. CSCWD
    EINS: Edge-Cloud Deep Model Inference with Network-Efficiency Schedule in Serverless
    Shijie Peng , Yanying Lin , Wenyan Chen, and 3 more authors
    In Proceedings of the 27th International Conference on Computer Supported Cooperative Work in Design, 2024