个人简介

我目前是北京大学人工智能研究院和集成电路学院双聘的预聘制助理教授。加入北京大学前,我曾任 Meta 端侧人工智能团队的高级研究科学家和技术负责人,从事面向下一代 AR/VR 设备的高效人工智能算法与硬件研究及产品化工作。我在David Z. Pan 教授的指导下获得德克萨斯大学奥斯汀分校计算机工程博士学位,并在黄如教授和王润声教授的指导下获得北京大学学士学位。

研究方向

我通过算法—硬件协同设计,研究高效且保护隐私的人工智能系统,并将其应用于大语言模型、具身智能和多模态智能。

🔥 诚招!

🔥 实验室有多个博士后岗位,研究方向为高效、安全AI加速。同时,我们欢迎有创造力、积极主动的实习生(本科生、硕士生)加入。

🔥 实验室每年有若干博士生和硕士生名额,优先考虑曾在组内实习的同学,欢迎尽早联系。

🔥 如有意向,请将简历和成绩单发送至邮件,邮件主题为**“Prospective Student from [Your Institute]”**。

下载我的个人简历。

兴趣爱好
  • 高效与安全的多模态人工智能
  • 算法/硬件协同设计与协同优化
教育经历
  • 博士,计算机工程, 2018

    德克萨斯大学奥斯汀分校,美国

  • 硕士,计算机工程, 2015

    德克萨斯大学奥斯汀分校,美国

  • 学士,微电子学, 2013

    北京大学,中国

工作经历

 
 
 
 
 
预聘制助理教授
2022年7月 – 现在 北京
人工智能研究院、集成电路学院
 
 
 
 
 
高级研究科学家
2018年9月 – 2022年7月 美国加利福尼亚州
  • 从研究科学家晋升为资深研究科学家、高级研究科学家
  • Meta Reality Labs 端侧人工智能技术负责人,聚焦面向 AR 眼镜的高效神经网络和软硬件协同设计
 
 
 
 
 
研究实习生(2016、2017年暑期)
2016年5月 – 2017年8月 美国加利福尼亚州
硬件安全与隐私保护人工智能
 
 
 
 
 
研发实习生
Cadence Design Systems
2014年5月 – 2014年8月 美国加利福尼亚州
静态时序分析加速

新闻动态

近期论文

完整论文列表请访问 Google Scholar

*
(2026). MatMoE: Matryoshka Mixture-of-Experts with Dynamic Mixed-Precision Quantization for Efficient Inference. In International Conference on Computer-Aided Design (ICCAD) 2026.

(2026). OptiPrime: Optimizing Private Inference through protocol-hardware codesign. In MICRO 2026.

(2026). Aging Aware Adaptive Voltage Scaling for Reliable and Efficient AI Accelerators. In International Symposium of Electronic Design Automation (ISEDA) 2026.

论文

(2026). NICE: 3D-NAND-based In-Memory-Computing with In-Situ ECC-Protection for Fault-Tolerant and Efficient LLM Inference. In International Symposium of Electronic Design Automation (ISEDA) 2026.

(2026). RePart: Efficient Hypergraph Partitioning with Logic Replication Optimization for Multi-FPGA System. In International Symposium of Electronic Design Automation (ISEDA) 2026.

论文

(2026). Attention Sink Forges Native MoE in Attention Layers: Sink-Aware Training to Address Head Collapse. In International Conference on Machine Learning (ICML) 2026.

论文

(2026). TEAM: Temporal–Spatial Consistency Guided Expert Activation for MoE Diffusion Language Model Acceleration. In International Conference on Machine Learning (ICML) 2026.

论文

(2026). Breaking the Reward Barrier: Accelerating Tree-of-Thought Reasoning via Speculative Exploration. In USENIX Symposium on Operating Systems Design and Implementation (OSDI) 2026.

论文

(2026). CREATE: Cross-Layer Resilience Characterization and Optimization for Efficient yet Reliable Embodied AI Systems. In International Conference on Architectural Support for Programming Languages and Operating Systems (ASPLOS) 2026.

论文

(2026). DRIFT: Harnessing Inherent Fault Tolerance for Efficient and Reliable Diffusion Model Inference. In Design Automation Conference (DAC) 2026.

论文

(2026). DySL-VLA: Efficient Vision-Language-Action Model Inference via Dynamic-Static Layer-Skipping for Robot Manipulation. In Design Automation Conference (DAC) 2026.

论文

(2026). EdgeSC: Universal Stochastic Computing Architecture for Efficient Edge Detection. In Design Automation Conference (DAC) 2026.

论文

(2026). KEEP: A KV-Cache-Centric Memory Management System for Efficient Embodied Planning. In Design Automation Conference (DAC) 2026.

论文

(2026). NASiC: 3D NAND-based CAM-Selected Multibit CIM Architecture for Efficient On-Device Mixture-of-Experts LLM Inference. In Design Automation Conference (DAC) 2026.

论文

(2026). Orchestrating Dual-Boundaries: An Arithmetic Intensity Inspired Acceleration Framework for Diffusion Language Models. In Design Automation Conference (DAC) 2026.

论文 代码

(2026). S2CIM: A Secure-Computation and Secure-Storage Compute-in-Memory Architecture with Circuit-Algorithm Co-Design for Efficient and Trustworthy Edge Inference. In Design Automation Conference (DAC) 2026.

论文

(2025). CryptoMoE: Privacy-Preserving and Scalable Mixture of Experts Inference via Balanced Expert Routing. In Conference on Neural Information Processing Systems (NeurIPs) 2025.

论文

(2025). EfficientNav: Towards On-Device Object-Goal Navigation with Navigation Map Caching and Retrieval. In Conference on Neural Information Processing Systems (NeurIPs) 2025.

论文

(2025). MPCache: MPC-Friendly KV Cache Eviction for Efficient Private LLM Inference. In Conference on Neural Information Processing Systems (NeurIPs) 2025.

论文

(2025). FENIX: Flexible and Efficient Hybrid HE/MPC Acceleration with Near-Memory Processing. In International Conference on Computer-Aided Design (ICCAD) 2025.

论文

(2025). H2EAL: Hybrid-Bonding Architecture with Hybrid Sparse Attention for Efficient Long-Context LLM Inference. In International Conference on Computer-Aided Design (ICCAD) 2025.

论文

(2025). HD-MoE: Hybrid and Dynamic Parallelism for Mixture-of-Expert LLMs with 3D Near-Memory Processing. In International Conference on Computer-Aided Design (ICCAD) 2025.

论文

(2025). No Redundancy, No Stall: Lightweight Streaming 3D Gaussian Splatting for Real-time Rendering. In International Conference on Computer-Aided Design (ICCAD) 2025.

论文

(2025). SpecMamba: Accelerating Mamba Inference on FPGA with Speculative Decoding. In International Conference on Computer-Aided Design (ICCAD) 2025.

论文

(2025). Ironman: Accelerating Oblivious Transfer Extension for Privacy-Preserving AI with Near-Memory Processing. In MICRO 2025.

论文

(2025). A 28nm 534.6TOPS/W Mixed-Precision Edge Accelerator for Embodied AI Using Stochastic Computing. In Asian Solid-State Circuits Conference (A-SSCC) 2025.

论文

(2025). HybriMoE: Hybrid CPU-GPU Scheduling and Cache Management for Efficient MoE Inference. In Design Automation Conference (DAC) 2025.

论文 项目 演示文稿

(2025). ReaLM: Reliable and Efficient Large Language Model Inference with Statistical Algorithm-Based Fault Tolerance. In Design Automation Conference (DAC) 2025.

论文 项目 演示文稿

(2025). Compact Non-Volatile Lookup Table Architecture based on Ferroelectric FET Array through In-Situ Combinatorial One-Hot Encoding for Reconfigurable Computing. In Design, Automation and Test in Europe Conference and Exhibition (DATE) 2025.

论文 演示文稿

(2025). FLASH: An Efficient Hardware Accelerator Leveraging Approximate and Sparse FFT for Homomorphic Encryption. In Design, Automation and Test in Europe Conference and Exhibition (DATE) 2025.

论文 演示文稿

(2025). LightMamba: Efficient Mamba Acceleration on FPGA with Quantization and Hardware Co-design. In Design, Automation and Test in Europe Conference and Exhibition (DATE) 2025.

论文 演示文稿

(2025). SCALES: Boost Binary Neural Network for Image Super-Resolution with Efficient Scalings. In Design, Automation and Test in Europe Conference and Exhibition (DATE) 2025.

论文 演示文稿

(2025). Stochastic Multivariate Universal-Radix Finite-State Machine: a Theoretically and Practically Elegant Nonlinear Function Approximator. In Asia and South Pacific Design Automation Conference (ASP-DAC) 2025.

论文

(2024). ArkVale: Efficient Generative LLM Inference with Recallable Key-Value Eviction. In Conference on Neural Information Processing Systems (NeurIPs) 2024.

论文

(2024). PrivCirNet: Efficient Private Inference via Block Circulant Transformation. In Conference on Neural Information Processing Systems (NeurIPs) 2024.

论文

(2024). AdapMoE: Adaptive Sensitivity-based Expert Gating and Management for Efficient MoE Inference. In International Conference on Computer-Aided Design (ICCAD) 2024.

论文

(2024). FlexHE: A flexible Kernel Generation Framework for Homomorphic Encryption-Based Private Inference. In International Conference on Computer-Aided Design (ICCAD) 2024.

论文

(2024). HG-PIPE: Vision Transformer Acceleration with Hybrid-Grained Pipeline. In International Conference on Computer-Aided Design (ICCAD) 2024.

论文

(2024). MCUBERT: Memory-Efficient BERT Inference on Commodity Microcontrollers. In International Conference on Computer-Aided Design (ICCAD) 2024.

论文

(2024). OSCA: End-to-end Serial Stochastic Computing Neural Acceleration with Fine-grained Scaling and Piecewise Activation. In International Conference on Computer-Aided Design (ICCAD) 2024.

论文

(2024). PrivQuant: Communication-Efficient Private Inference with Quantized Network/Protocol Co-Optimization. In International Conference on Computer-Aided Design (ICCAD) 2024.

论文

(2024). ProPD: Dynamic Token Tree Pruning and Generation for LLM Parallel Decoding. In International Conference on Computer-Aided Design (ICCAD) 2024.

论文

(2024). CASCADE: A Framework for CNN Accelerator Synthesis with Concatenation and Refreshing Dataflow. In IEEE Transactions on Circuits and Systems I: Regular Papers (TCAS-I) (2024).

论文

(2024). FastQuery: Communication-efficient Embedding Table Query for Private LLMs inference. In Design Automation Conference (DAC) 2024.

论文

(2024). ASCEND: Accurate yet Efficient End-to-End Stochastic Computing Acceleration of Vision Transformer. In Design, Automation and Test in Europe Conference and Exhibition (DATE) 2024.

论文

(2024). A 16.38TOPS and 4.55POPS/W SRAM Computing-in-Memory Macro for Signed Operands Computation and Batch Normalization Implementation. In IEEE Transactions on Circuits and Systems I: Regular Papers (TCAS-I) (2024).

论文

(2024). MixCIM: A Hybrid-Cell-Based Computing-in-Memory Macro with Less-Data-Movement and Activation-Memory-Reuse for Depthwise Separable Neural Networks. In IEEE Custom Integrated Circuits Conference (CICC) 2024.

论文

(2023). CoPriv: Network/Protocol Co-Optimization for Communication-Efficient Private Inference. In Conference on Neural Information Processing Systems (NeurIPs) 2023.

论文

(2023). Falcon: Accelerating Homomorphically Encrypted Convolutions for Efficient Private Mobile Network Inference. In ACM/IEEE International Conference on Computer Aided Design (ICCAD) 2023.

论文

(2023). Memory-aware Scheduling for Complex Wired Networks with Iterative Graph Optimization. In ACM/IEEE International Conference on Computer Aided Design (ICCAD) 2023.

论文

(2023). MPCViT: Searching for Accurate and Efficient MPC-Friendly Vision Transformer with Heterogeneous Attention. In International Conference on Computer Vision (ICCV) 2023.

论文

(2023). Not your father’s stochastic computing (SC)! Efficient yet Accurate End-to-End SC Accelerator Design. In International Conference on ASIC (ASICON) 2023.

论文

(2023). READ: Reliability-Enhanced Accelerator Dataflow Optimization using Critical Input Pattern Reduction. In ACM/IEEE International Conference on Computer Aided Design (ICCAD) 2023.

论文

(2023). AVATAR: An Aging- and Variation-Aware Dynamic Timing Analyzer for Error-Efficient Computing. In IEEE Transactions on Computer-Aided Design of Integrated Circuits and Systems (TCAD) (2023).

论文

(2023). Efficient Non-Linear Adder for Stochastic Computing with Approximate Spatial-Temporal Sorting Network. In Design Automation Conference (DAC) 2023.

论文

(2023). Accurate yet Efficient Stochastic Computing Neural Acceleration with High Precision Residual Fusion. In Design, Automation and Test in Europe Conference and Exhibition (DATE) 2023.

论文

(2023). READ: Reliability-Enhanced Accelerator Dataflow Optimization using Critical Input Pattern Reduction. In Design, Automation and Test in Europe Conference and Exhibition (DATE) 2023 (extended abstract).

论文

(2022). BiT: Robustly Binarized Multi-distilled Transformer. In Conference on Neural Information Processing Systems (NeurIPs) 2022.

论文

(2022). Depth Shrink: Empowering Hardware-Friendly Shallow Neural Networks. In Conference on Machine Learning (ICML) 2022.

(2022). Omni-sparsity DNN: Fast Sparsity Optimization for On-Device Streaming E2E ASR via Supernet. In International Conference on Acoustics, Speech and Signal Processing (ICASSP) 2022.

论文

(2022). Multi-Scale High-Resolution Vision Transformer for Semantic Segmentation. In Conference on Computer Vision and Pattern Recognition (CVPR) 2022.

论文

(2022). SplitNets: Designing Neural Architectures for Efficient Distributed Computing on Head-Mounted Systems. In Conference on Computer Vision and Pattern Recognition (CVPR) 2022.

论文

(2022). NASViT: Neural Architecture Search for Efficient Vision Transformers with Gradient Conflict aware Supernet Training. In Conference on Learning Representations (ICLR) 2022.

论文

(2021). DNA: Differentiable Network-Accelerator Co-Search. In International Symposium on Low Power Electronics and Design (ISLPED) 2021.

论文

(2021). AlphaNet: Improved Training of Supernets with Alpha-Divergence. In Conference on Machine Learning (ICML) 2021 (Long Oral).

论文

(2021). AttentiveNAS: Improving Neural Architecture Search via Attentive Sampling. In Conference on Computer Vision and Pattern Recognition (CVPR) 2021.

论文

(2021). Improving efficiency in neural network accelerator using operands hamming distance optimization. In Asia and South Pacific Design Automation Conference (ASP-DAC) 2021.

论文

(2020). KeepAugment: A Simple Information-Preserving Data Augmentation Approach. In Conference on Computer Vision and Pattern Recognition (CVPR) 2021.

论文

(2020). Co-exploration of neural architectures and heterogeneous asic accelerator designs targeting multiple tasks. In ACM/IEEE Design Automation Conference (DAC) 2020.

论文

(2018). TimingSAT: Decamouflaging timing-based logic obfuscation. In IEEE International Test Conference (ITC) 2018.

论文

(2018). A Practical Split Manufacturing Framework for Trojan Prevention via Simultaneous Wire Lifting and Cell Insertion. In IEEE Transactions on Computer-Aided Design of Integrated Circuits and Systems (TCAD) (2018).

论文

(2018). Federated Learning with Non-IID Data. In arXiv:1806.00582 (2018).

论文

(2018). A Practical Split Manufacturing Framework for Trojan Prevention via Simultaneous Wire Lifting and Cell Insertion. In Asia and South Pacific Design Automation Conference (ASP-DAC) 2018.

论文

(2018). PrivyNet: A Flexible Framework for Privacy-Preserving Deep Neural Network Training. In arXiv:1709:06161 (2018).

论文

(2017). Provably secure camouflaging strategy for IC protection. In ACM/IEEE International Conference on Computer Aided Design (ICCAD) 2017.

论文

(2017). Provably secure camouflaging strategy for IC protection. In IEEE Transactions on Computer-Aided Design of Integrated Circuits and Systems (TCAD) (2018).

论文

(2017). Cross-level monte carlo framework for system vulnerability evaluation against fault attack. In ACM/IEEE Design Automation Conference (DAC) 2017.

论文

(2017). AppSAT: Approximately Deobfuscating Integrated Circuits. In IEEE International Symposium on Hardware Oriented Security and Trust (HOST) 2017 (Best Paper Award).

论文

(2016). A monte carlo simulation flow for seu analysis of sequential circuits. In ACM/IEEE Design Automation Conference (DAC) 2016.

论文

(2016). Practical public PUF enabled by solving max-flow problem on chip. In ACM/IEEE Design Automation Conference (DAC) 2016.

论文

(2013). Characterization of Random Telegraph Noise in Scaled High-κ/Metal-Gate MOSFETs with SiO2/HfO2 Gate Dielectrics. In ECS Transactions, 52 (1) 941-946 (2013).

论文

(0001). Alchemist: A Unified Accelerator Architecture for Cross-Scheme Fully Homomorphic Encryption.

(0001). Enhancing 3D Detection Through Feature Aligned Deep Fusion.

(0001). MoteNN: Memory Optimization via Fine-grained Scheduling for Deep Neural Networks on Tiny Devices.

(0001). Swift: Fast Secure Neural Network Inference with Fully Homomorphic Encryption.

开源项目

更多开源项目请访问 GitHub

*

荣誉与奖励

  • 2026年北京大学教学优秀奖 — 北京大学, 2026年8月
  • ICML Gold Reviewer — ICML, 2026年5月
  • ACM SIGDA 杰出青年教师奖 — ACM SIGDA, 2026年4月
  • 蚂蚁集团 InTech 未来奖 — 蚂蚁集团, 2025年9月
  • 端侧多模态生成式科学智能竞赛第一名 — ACM SIGDA, 2025年6月
  • AICAS 大语言模型硬件系统设计挑战赛第一名 — IEEE Circuits and Systems Society, 2025年4月
  • 青年教师教学基本功比赛一等奖 — 北京大学, 2024年12月
  • CCF—蚂蚁科研基金软硬件协同设计专项 — 中国计算机学会(CCF), 2024年8月
  • CCF 集成电路 Early Career Award — 中国计算机学会(CCF), 2024年7月
  • 隐语产学合作杰出贡献奖 — 蚂蚁集团, 2024年5月
  • CCF—蚂蚁科研基金隐私计算专项 — 中国计算机学会(CCF), 2023年8月
  • Margarida Jacome 杰出博士论文奖 — 德克萨斯大学奥斯汀分校, 2019年5月
  • 杰出博士论文奖 — 欧洲设计自动化协会(EDAA), 2019年5月
  • ACM 博士学位论文奖提名 — 德克萨斯大学奥斯汀分校, 2019年3月
  • 学生科研竞赛全球总决赛研究生组第一名 — Association for Computing Machinery (ACM), 2018年6月
  • 最佳论文奖 — ACM Great Lakes Symposium on VLSI (GLSVLSI), 2018年3月
  • 最佳海报展示奖 — ASP-DAC Student Research Forum, ACM SIGDA, 2018年2月
  • ICCAD博士生科研竞赛金牌 — ACM SIGDA, 2017年11月
  • 最佳论文奖 — IEEE International Symposium on Hardware Oriented Security and Trust (HOST), 2017年7月
  • Cockrell 工程学院研究生奖学金 — 德克萨斯大学奥斯汀分校, 2013年9月
  • 杨芙清—王阳元院士奖学金 — 北京大学, 2011年9月
  • 李彦宏百度奖学金 — 北京大学, 2010年9月

联系方式