Research

Our research ranges from core microarchitectures to multi-node AI systems. Across these environments, we jointly optimize algorithms, system software, and physical hardware as a single integrated design space to maximize end-to-end efficiency.

Hardware-Software Co-Design for AI

We co-design domain-specific hardware accelerators and hardware-aware algorithms for efficient AI acceleration. By jointly optimizing neural network quantization, sparsity (pruning), and datapath architectures, we maximize the end-to-end efficiency of AI workloads. Through this holistic approach, we aim to bridge the gap between diverse software algorithms and underlying hardware structures.

On-Device AI Acceleration

We develop hardware-aware execution frameworks to deploy massive multimodal AI models on power-constrained edge devices. Executing complex workloads, such as Vision-Language-Action (VLA) models, requires tailoring advanced model compression and system-level optimizations to specific edge processors. We map these optimized workloads onto actual edge hardware to profile their real-world latency, memory footprint, and power consumption. This practical evaluation ensures that our diverse optimization strategies translate to viable on-device AI solutions.

System Architecture for AI

We design hardware architectures and optimize system stacks to accelerate large-scale AI across multiple compute nodes. To overcome the interconnect and memory bottlenecks of datacenter AI, we develop scalable systems ranging from Processing-Near-Memory (PNM) hardware to efficient multi-GPU configurations. By leveraging technologies such as Compute Express Link (CXL) and refining data orchestration across nodes, we enable scalable execution of large-scale AI models.

Hardware Microarchitecture

We design and evaluate core microarchitectures and memory systems to advance the fundamental performance of modern processors. We investigate essential hardware structures that determine overall processing efficiency, focusing on bandwidth-efficient memory hierarchies, specialized execution units, and optimized data movement. By building detailed cycle-accurate simulators, we quantitatively analyze new microarchitectural mechanisms that maximize compute density and hardware utilization.