Shuzhang Zhong

PhD Student (2023)

Peking University

Interests
  • Efficient AI
Education
  • B.S. in Computer Science And Technology, 2023

    Beihang University

Publications

(2026). Breaking the Reward Barrier: Accelerating Tree-of-Thought Reasoning via Speculative Exploration. In USENIX Symposium on Operating Systems Design and Implementation (OSDI) 2026.

Paper

(2026). NASiC: 3D NAND-based CAM-Selected Multibit CIM Architecture for Efficient On-Device Mixture-of-Experts LLM Inference. In Design Automation Conference (DAC) 2026.

Paper

(2025). H2EAL: Hybrid-Bonding Architecture with Hybrid Sparse Attention for Efficient Long-Context LLM Inference. In International Conference on Computer-Aided Design (ICCAD) 2025.

Paper

(2025). HD-MoE: Hybrid and Dynamic Parallelism for Mixture-of-Expert LLMs with 3D Near-Memory Processing. In International Conference on Computer-Aided Design (ICCAD) 2025.

Paper

(2025). HybriMoE: Hybrid CPU-GPU Scheduling and Cache Management for Efficient MoE Inference. In Design Automation Conference (DAC) 2025.

Paper Project Slides

(2025). SpecASR: Accelerating LLM-based Automatic Speech Recognition via Speculative Decoding. In Design Automation Conference (DAC) 2025.

Paper Slides

(2024). AdapMoE: Adaptive Sensitivity-based Expert Gating and Management for Efficient MoE Inference. In International Conference on Computer-Aided Design (ICCAD) 2024.

Paper

(2024). PrivQuant: Communication-Efficient Private Inference with Quantized Network/Protocol Co-Optimization. In International Conference on Computer-Aided Design (ICCAD) 2024.

Paper

(2024). ProPD: Dynamic Token Tree Pruning and Generation for LLM Parallel Decoding. In International Conference on Computer-Aided Design (ICCAD) 2024.

Paper

(2023). Memory-aware Scheduling for Complex Wired Networks with Iterative Graph Optimization. In ACM/IEEE International Conference on Computer Aided Design (ICCAD) 2023.

Paper