Biography

I am currently a tenure-track assistant professor jointly affiliated with the Institute for Artificial Intelligence and the School of Integrated Circuits at Peking University. Before joining Peking University, I was a staff research scientist and tech lead on Meta’s On-Device AI team, where I researched and productized efficient AI algorithms and hardware for next-generation AR/VR devices. I received my Ph.D. in Computer Engineering from the University of Texas at Austin under the supervision of Prof. David Z. Pan and my bachelor’s degree from Peking University under the supervision of Prof. Ru Huang and Prof. Runsheng Wang.

Research Focus

I develop efficient and privacy-preserving AI systems through algorithm–hardware co-design, with applications in large language models, embodied AI, and multimodal intelligence.

🔥 We are actively recruiting!

🔥 My group has multiple Postdoctoral positions available with a focus on Efficient and Secure AI Acceleration. We also welcome creative and self-motivated interns (undergraduate and Master’s students).

🔥 My group often has several PhD and master positions each year. We always prioritize students interning in the group. Please contact early.

🔥 If you are interested, please send me an email with subject “Prospective Student from [Your Institute]” along with your CV and transcripts.

Download my resumé.

Interests
  • Efficient and Secure Multi-Modality Artificial Intelligence
  • Algorithm/Hardware Co-Design/Co-Optimization
Education
  • PhD in Computer Engineering, 2018

    University of Texas at Austin, Austin, Tx, USA

  • MS in Computer Engineering, 2015

    University of Texas at Austin, Austin, Tx, USA

  • BS in Microelectronics, 2013

    Peking University, Beijing, China

Experience

 
 
 
 
 
Tenure-Track Assistant Professor
Jul 2022 – Present Beijing
Institute for Artificial Intelligence & School of Integrated Circuits
 
 
 
 
 
Staff Research Scientist
Sep 2018 – Jul 2022 California
  • Progressed from Research Scientist to Senior and Staff Research Scientist
  • Tech Lead for on-device AI at Meta Reality Labs, focusing on efficient neural networks and hardware co-design for AR glasses
 
 
 
 
 
Research Intern (Summer 2016 & 2017)
May 2016 – Aug 2017 California
Hardware security and privacy-preserving AI
 
 
 
 
 
Research & Design Intern
Cadence Design System
May 2014 – Aug 2014 California
Static timing analysis acceleration

Recent Publications

More detailed publication lists available through Google Scholar

*
(2026). MatMoE: Matryoshka Mixture-of-Experts with Dynamic Mixed-Precision Quantization for Efficient Inference. In International Conference on Computer-Aided Design (ICCAD) 2026.

(2026). OptiPrime: Optimizing Private Inference through protocol-hardware codesign. In MICRO 2026.

(2026). Aging Aware Adaptive Voltage Scaling for Reliable and Efficient AI Accelerators. In International Symposium of Electronic Design Automation (ISEDA) 2026.

Paper

(2026). NICE: 3D-NAND-based In-Memory-Computing with In-Situ ECC-Protection for Fault-Tolerant and Efficient LLM Inference. In International Symposium of Electronic Design Automation (ISEDA) 2026.

(2026). RePart: Efficient Hypergraph Partitioning with Logic Replication Optimization for Multi-FPGA System. In International Symposium of Electronic Design Automation (ISEDA) 2026.

Paper

(2026). Attention Sink Forges Native MoE in Attention Layers: Sink-Aware Training to Address Head Collapse. In International Conference on Machine Learning (ICML) 2026.

Paper

(2026). TEAM: Temporal–Spatial Consistency Guided Expert Activation for MoE Diffusion Language Model Acceleration. In International Conference on Machine Learning (ICML) 2026.

Paper

(2026). Breaking the Reward Barrier: Accelerating Tree-of-Thought Reasoning via Speculative Exploration. In USENIX Symposium on Operating Systems Design and Implementation (OSDI) 2026.

Paper

(2026). CREATE: Cross-Layer Resilience Characterization and Optimization for Efficient yet Reliable Embodied AI Systems. In International Conference on Architectural Support for Programming Languages and Operating Systems (ASPLOS) 2026.

Paper

(2026). DRIFT: Harnessing Inherent Fault Tolerance for Efficient and Reliable Diffusion Model Inference. In Design Automation Conference (DAC) 2026.

Paper

(2026). DySL-VLA: Efficient Vision-Language-Action Model Inference via Dynamic-Static Layer-Skipping for Robot Manipulation. In Design Automation Conference (DAC) 2026.

Paper

(2026). EdgeSC: Universal Stochastic Computing Architecture for Efficient Edge Detection. In Design Automation Conference (DAC) 2026.

Paper

(2026). KEEP: A KV-Cache-Centric Memory Management System for Efficient Embodied Planning. In Design Automation Conference (DAC) 2026.

Paper

(2026). NASiC: 3D NAND-based CAM-Selected Multibit CIM Architecture for Efficient On-Device Mixture-of-Experts LLM Inference. In Design Automation Conference (DAC) 2026.

Paper

(2026). Orchestrating Dual-Boundaries: An Arithmetic Intensity Inspired Acceleration Framework for Diffusion Language Models. In Design Automation Conference (DAC) 2026.

Paper Code

(2026). S2CIM: A Secure-Computation and Secure-Storage Compute-in-Memory Architecture with Circuit-Algorithm Co-Design for Efficient and Trustworthy Edge Inference. In Design Automation Conference (DAC) 2026.

Paper

(2025). CryptoMoE: Privacy-Preserving and Scalable Mixture of Experts Inference via Balanced Expert Routing. In Conference on Neural Information Processing Systems (NeurIPs) 2025.

Paper

(2025). EfficientNav: Towards On-Device Object-Goal Navigation with Navigation Map Caching and Retrieval. In Conference on Neural Information Processing Systems (NeurIPs) 2025.

Paper

(2025). MPCache: MPC-Friendly KV Cache Eviction for Efficient Private LLM Inference. In Conference on Neural Information Processing Systems (NeurIPs) 2025.

Paper

(2025). FENIX: Flexible and Efficient Hybrid HE/MPC Acceleration with Near-Memory Processing. In International Conference on Computer-Aided Design (ICCAD) 2025.

Paper

(2025). H2EAL: Hybrid-Bonding Architecture with Hybrid Sparse Attention for Efficient Long-Context LLM Inference. In International Conference on Computer-Aided Design (ICCAD) 2025.

Paper

(2025). HD-MoE: Hybrid and Dynamic Parallelism for Mixture-of-Expert LLMs with 3D Near-Memory Processing. In International Conference on Computer-Aided Design (ICCAD) 2025.

Paper

(2025). No Redundancy, No Stall: Lightweight Streaming 3D Gaussian Splatting for Real-time Rendering. In International Conference on Computer-Aided Design (ICCAD) 2025.

Paper

(2025). SpecMamba: Accelerating Mamba Inference on FPGA with Speculative Decoding. In International Conference on Computer-Aided Design (ICCAD) 2025.

Paper

(2025). Ironman: Accelerating Oblivious Transfer Extension for Privacy-Preserving AI with Near-Memory Processing. In MICRO 2025.

Paper

(2025). A 28nm 534.6TOPS/W Mixed-Precision Edge Accelerator for Embodied AI Using Stochastic Computing. In Asian Solid-State Circuits Conference (A-SSCC) 2025.

Paper

(2025). HybriMoE: Hybrid CPU-GPU Scheduling and Cache Management for Efficient MoE Inference. In Design Automation Conference (DAC) 2025.

Paper Project Slides

(2025). ReaLM: Reliable and Efficient Large Language Model Inference with Statistical Algorithm-Based Fault Tolerance. In Design Automation Conference (DAC) 2025.

Paper Project Slides

(2025). SpecASR: Accelerating LLM-based Automatic Speech Recognition via Speculative Decoding. In Design Automation Conference (DAC) 2025.

Paper Slides

(2025). Compact Non-Volatile Lookup Table Architecture based on Ferroelectric FET Array through In-Situ Combinatorial One-Hot Encoding for Reconfigurable Computing. In Design, Automation and Test in Europe Conference and Exhibition (DATE) 2025.

Paper Slides

(2025). FLASH: An Efficient Hardware Accelerator Leveraging Approximate and Sparse FFT for Homomorphic Encryption. In Design, Automation and Test in Europe Conference and Exhibition (DATE) 2025.

Paper Slides

(2025). LightMamba: Efficient Mamba Acceleration on FPGA with Quantization and Hardware Co-design. In Design, Automation and Test in Europe Conference and Exhibition (DATE) 2025.

Paper Slides

(2025). SCALES: Boost Binary Neural Network for Image Super-Resolution with Efficient Scalings. In Design, Automation and Test in Europe Conference and Exhibition (DATE) 2025.

Paper Slides

(2025). Stochastic Multivariate Universal-Radix Finite-State Machine: a Theoretically and Practically Elegant Nonlinear Function Approximator. In Asia and South Pacific Design Automation Conference (ASP-DAC) 2025.

Paper

(2024). ArkVale: Efficient Generative LLM Inference with Recallable Key-Value Eviction. In Conference on Neural Information Processing Systems (NeurIPs) 2024.

Paper

(2024). PrivCirNet: Efficient Private Inference via Block Circulant Transformation. In Conference on Neural Information Processing Systems (NeurIPs) 2024.

Paper

(2024). AdapMoE: Adaptive Sensitivity-based Expert Gating and Management for Efficient MoE Inference. In International Conference on Computer-Aided Design (ICCAD) 2024.

Paper

(2024). FlexHE: A flexible Kernel Generation Framework for Homomorphic Encryption-Based Private Inference. In International Conference on Computer-Aided Design (ICCAD) 2024.

Paper

(2024). HG-PIPE: Vision Transformer Acceleration with Hybrid-Grained Pipeline. In International Conference on Computer-Aided Design (ICCAD) 2024.

Paper

(2024). MCUBERT: Memory-Efficient BERT Inference on Commodity Microcontrollers. In International Conference on Computer-Aided Design (ICCAD) 2024.

Paper

(2024). OSCA: End-to-end Serial Stochastic Computing Neural Acceleration with Fine-grained Scaling and Piecewise Activation. In International Conference on Computer-Aided Design (ICCAD) 2024.

Paper

(2024). PrivQuant: Communication-Efficient Private Inference with Quantized Network/Protocol Co-Optimization. In International Conference on Computer-Aided Design (ICCAD) 2024.

Paper

(2024). ProPD: Dynamic Token Tree Pruning and Generation for LLM Parallel Decoding. In International Conference on Computer-Aided Design (ICCAD) 2024.

Paper

(2024). CASCADE: A Framework for CNN Accelerator Synthesis with Concatenation and Refreshing Dataflow. In IEEE Transactions on Circuits and Systems I: Regular Papers (TCAS-I) (2024).

Paper

(2024). FastQuery: Communication-efficient Embedding Table Query for Private LLMs inference. In Design Automation Conference (DAC) 2024.

Paper

(2024). ASCEND: Accurate yet Efficient End-to-End Stochastic Computing Acceleration of Vision Transformer. In Design, Automation and Test in Europe Conference and Exhibition (DATE) 2024.

Paper

(2024). A 16.38TOPS and 4.55POPS/W SRAM Computing-in-Memory Macro for Signed Operands Computation and Batch Normalization Implementation. In IEEE Transactions on Circuits and Systems I: Regular Papers (TCAS-I) (2024).

Paper

(2024). MixCIM: A Hybrid-Cell-Based Computing-in-Memory Macro with Less-Data-Movement and Activation-Memory-Reuse for Depthwise Separable Neural Networks. In IEEE Custom Integrated Circuits Conference (CICC) 2024.

Paper

(2023). CoPriv: Network/Protocol Co-Optimization for Communication-Efficient Private Inference. In Conference on Neural Information Processing Systems (NeurIPs) 2023.

Paper

(2023). Falcon: Accelerating Homomorphically Encrypted Convolutions for Efficient Private Mobile Network Inference. In ACM/IEEE International Conference on Computer Aided Design (ICCAD) 2023.

Paper

(2023). Memory-aware Scheduling for Complex Wired Networks with Iterative Graph Optimization. In ACM/IEEE International Conference on Computer Aided Design (ICCAD) 2023.

Paper

(2023). MPCViT: Searching for Accurate and Efficient MPC-Friendly Vision Transformer with Heterogeneous Attention. In International Conference on Computer Vision (ICCV) 2023.

Paper

(2023). Not your father’s stochastic computing (SC)! Efficient yet Accurate End-to-End SC Accelerator Design. In International Conference on ASIC (ASICON) 2023.

Paper

(2023). READ: Reliability-Enhanced Accelerator Dataflow Optimization using Critical Input Pattern Reduction. In ACM/IEEE International Conference on Computer Aided Design (ICCAD) 2023.

Paper

(2023). AVATAR: An Aging- and Variation-Aware Dynamic Timing Analyzer for Error-Efficient Computing. In IEEE Transactions on Computer-Aided Design of Integrated Circuits and Systems (TCAD) (2023).

Paper

(2023). Efficient Non-Linear Adder for Stochastic Computing with Approximate Spatial-Temporal Sorting Network. In Design Automation Conference (DAC) 2023.

Paper

(2023). Accurate yet Efficient Stochastic Computing Neural Acceleration with High Precision Residual Fusion. In Design, Automation and Test in Europe Conference and Exhibition (DATE) 2023.

Paper

(2023). READ: Reliability-Enhanced Accelerator Dataflow Optimization using Critical Input Pattern Reduction. In Design, Automation and Test in Europe Conference and Exhibition (DATE) 2023 (extended abstract).

Paper

(2022). BiT: Robustly Binarized Multi-distilled Transformer. In Conference on Neural Information Processing Systems (NeurIPs) 2022.

Paper

(2022). Depth Shrink: Empowering Hardware-Friendly Shallow Neural Networks. In Conference on Machine Learning (ICML) 2022.

(2022). Omni-sparsity DNN: Fast Sparsity Optimization for On-Device Streaming E2E ASR via Supernet. In International Conference on Acoustics, Speech and Signal Processing (ICASSP) 2022.

Paper

(2022). Multi-Scale High-Resolution Vision Transformer for Semantic Segmentation. In Conference on Computer Vision and Pattern Recognition (CVPR) 2022.

Paper

(2022). SplitNets: Designing Neural Architectures for Efficient Distributed Computing on Head-Mounted Systems. In Conference on Computer Vision and Pattern Recognition (CVPR) 2022.

Paper

(2022). NASViT: Neural Architecture Search for Efficient Vision Transformers with Gradient Conflict aware Supernet Training. In Conference on Learning Representations (ICLR) 2022.

Paper

(2021). DNA: Differentiable Network-Accelerator Co-Search. In International Symposium on Low Power Electronics and Design (ISLPED) 2021.

Paper

(2021). AlphaNet: Improved Training of Supernets with Alpha-Divergence. In Conference on Machine Learning (ICML) 2021 (Long Oral).

Paper

(2021). AttentiveNAS: Improving Neural Architecture Search via Attentive Sampling. In Conference on Computer Vision and Pattern Recognition (CVPR) 2021.

Paper

(2021). Improving efficiency in neural network accelerator using operands hamming distance optimization. In Asia and South Pacific Design Automation Conference (ASP-DAC) 2021.

Paper

(2020). KeepAugment: A Simple Information-Preserving Data Augmentation Approach. In Conference on Computer Vision and Pattern Recognition (CVPR) 2021.

Paper

(2020). Co-exploration of neural architectures and heterogeneous asic accelerator designs targeting multiple tasks. In ACM/IEEE Design Automation Conference (DAC) 2020.

Paper

(2018). TimingSAT: Decamouflaging timing-based logic obfuscation. In IEEE International Test Conference (ITC) 2018.

Paper

(2018). A Synergistic Framework for Hardware IP Privacy and Integrity Protection. In Springer (2018).

Paper

(2018). A Practical Split Manufacturing Framework for Trojan Prevention via Simultaneous Wire Lifting and Cell Insertion. In IEEE Transactions on Computer-Aided Design of Integrated Circuits and Systems (TCAD) (2018).

Paper

(2018). Federated Learning with Non-IID Data. In arXiv:1806.00582 (2018).

Paper

(2018). A Practical Split Manufacturing Framework for Trojan Prevention via Simultaneous Wire Lifting and Cell Insertion. In Asia and South Pacific Design Automation Conference (ASP-DAC) 2018.

Paper

(2018). PrivyNet: A Flexible Framework for Privacy-Preserving Deep Neural Network Training. In arXiv:1709:06161 (2018).

Paper

(2017). Provably secure camouflaging strategy for IC protection. In ACM/IEEE International Conference on Computer Aided Design (ICCAD) 2017.

Paper

(2017). Provably secure camouflaging strategy for IC protection. In IEEE Transactions on Computer-Aided Design of Integrated Circuits and Systems (TCAD) (2018).

Paper

(2017). Cross-level monte carlo framework for system vulnerability evaluation against fault attack. In ACM/IEEE Design Automation Conference (DAC) 2017.

Paper

(2017). AppSAT: Approximately Deobfuscating Integrated Circuits. In IEEE International Symposium on Hardware Oriented Security and Trust (HOST) 2017 (Best Paper Award).

Paper

(2016). A monte carlo simulation flow for seu analysis of sequential circuits. In ACM/IEEE Design Automation Conference (DAC) 2016.

Paper

(2016). Practical public PUF enabled by solving max-flow problem on chip. In ACM/IEEE Design Automation Conference (DAC) 2016.

Paper

(2013). Characterization of Random Telegraph Noise in Scaled High-κ/Metal-Gate MOSFETs with SiO2/HfO2 Gate Dielectrics. In ECS Transactions, 52 (1) 941-946 (2013).

Paper

(0001). Alchemist: A Unified Accelerator Architecture for Cross-Scheme Fully Homomorphic Encryption.

(0001). Enhancing 3D Detection Through Feature Aligned Deep Fusion.

(0001). MoteNN: Memory Optimization via Fine-grained Scheduling for Deep Neural Networks on Tiny Devices.

(0001). Swift: Fast Secure Neural Network Inference with Fully Homomorphic Encryption.

Projects

More opensource projects available through github

*

Accomplish­ments

  • ACM SIGDA Outstanding Young Faculty Award — ACM SIGDA, Apr 2026
  • Ant Group InTech Future Award — Ant Group, Sep 2025
  • On-Device Multi-modal Generative AI for Science Contest 1st Place — ACM SIGDA, Jun 2025
  • AICAS Grand Challenge on LLM Hardware System Design 1st Place — IEEE Circuit & System Society, Apr 2025
  • Young Teachers’ Teaching Skills Competition 1st Place Prize — Peking University, Dec 2024
  • CCF-Ant Group Research Award on Hardware/Software Co-Design — China Computer Federation (CCF), Aug 2024
  • CCF Integrated Circuits Early Career Award — China Computer Federation (CCF), Jul 2024
  • Secretflow Outstanding Industry-Academic Cooperation Contribution Award — Ant Group, May 2024
  • CCF-Ant Group Research Award on Privacy Computing — China Computer Federation (CCF), Aug 2023
  • Margarida Jacome Outstanding Dissertation Prize — University of Texas at Austin, May 2019
  • Outstanding Dissertations Award — European Design and Automation Association (EDAA), May 2019
  • Nominee of ACM Doctoral Dissertation Award — University of Texas at Austin, Mar 2019
  • First Place, Student Research Competition Grand Final (Graduate Category) — Aossication of Computing Machinery (ACM), Jun 2018
  • Best Paper Award — ACM Great Lake Symposium on VLSI (GLSVLSI), Mar 2018
  • Best Poster (Presentation) Award — ASPDAC Student Research Forum, Aossication of Computing Machinery (ACM) SIGDA, Feb 2018
  • Gold Medal — ICCAD Student Research Competition, Aossication of Computing Machinery (ACM) SIGDA, Nov 2017
  • Best Paper Award — IEEE International Symposium on Hardware Oriented Security and Trust (HOST), Jul 2017
  • Cockrell School Graduate Student Fellowship — University of Texas at Austin, Sep 2013
  • Yang Fuqing and Wang Yangyuan Academician Scholarship — Peking University, Sep 2011
  • Li Yanhong Baidu Scholarship — Peking University, Sep 2010

Contact