Publications
Listed in reverse chronological order.
2026
- KV Admission: Learning What to Write for Efficient Long-Context LLM InferenceYen-Chieh Huang, Pi-Cheng Hsiu, Rui Fang, and Ming-Syan ChenTo appear in Conference on Empirical Methods in Natural Language Processing (EMNLP), 2026
Long-context LLM inference is bottlenecked by the quadratic attention complexity and linear Key-Value (KV) cache growth. Prior approaches mitigate this via post-hoc selection or eviction but overlook the root inefficiency: indiscriminate token admission. In this paper, we formalize KV management as a causal system of three primitives: KV Admission, Selection, and Eviction. We instantiate KV Admission via Write-Gated KV (WG-KV), a lightweight mechanism that learns to predict token utility before cache entry. By filtering out redundant states early to maintain a compact global cache alongside a sliding local cache, WG-KV significantly reduces memory usage and accelerates both prefill and decode phases. Our results demonstrate that learning what to write is a principled and practical recipe for efficient long-context inference. Code is available at https://github.com/EMCLab-Sinica/WG-KV.
- Splitting Bottlenecks: Memory-Aware Neural Architecture Search for Multi-Branch TinyMLChia-Yin Liu, Hashan Roshantha Mendis, Yen-Chieh Huang, Yi-Jung Chen, and Pi-Cheng HsiuTo appear in IEEE Transactions on Computer-Aided Design of Integrated Circuits and Systems (TCAD), 2026
Bringing modern intelligence to tiny devices is challenging because severe on-chip memory limits fundamentally constrain the multi-branch architectures prevalent in modern neural networks. Tensor splitting is a common technique to reduce peak memory usage. However, existing splitting strategies designed for single-branch networks fail to effectively reduce memory from retained tensors in multi-branch networks, where operator scheduling dominates the peak and splitting effectiveness varies widely across architectures.We present DupNAS, a framework that jointly optimizes neural architecture and multi-branch tensor splitting to enable higher-accuracy architectures under tighter memory constraints. In multi-branch networks, combinatorial operator schedules induce a vast splitting configuration space, making direct integration with neural architecture search intractable. DupNAS addresses this challenge by aggressively shrinking the configuration space without sacrificing peak memory optimality, while incrementally exploring configurations that resolve memory bottlenecks. We evaluate DupNAS across diverse vision models and memory constraints, as well as deploy the resulting networks on STM32 microcontrollers. By substantially reducing peak memory, DupNAS improves accuracy by 5 percentage points on average over existing splitting strategies at comparable optimization time, while achieving superior on-device accuracy-latency trade-offs.
2025
- AQB8: Energy-Efficient Ray Tracing Accelerator through Multi-Level QuantizationYen-Chieh Huang, Chen-Pin Yang, and Tsung Tai YehIn International Symposium on Computer Architecture (ISCA), 2025
Ray tracing (RT) is a rendering technique that produces high-fidelity images by simulating how light physically interacts with objects in a scene. This realism comes at a high computational and memory cost, largely driven by the need to find numerous ray-object intersections. To accelerate this process, scenes are typically structured using bounding volume hierarchy (BVH) trees, where objects are grouped within bounding boxes to minimize the number of required intersection tests. Specialized hardware, known as RT accelerators, accelerates BVH processing to boost computational speed, yet memory traffic persists as a major bottleneck, largely due to the bandwidth consumed by standard 32-bit floating-point (FP32) bounding boxes. Consequently, previous work has aimed to reduce memory traffic by compressing these boxes into low-bit (e.g., 8-bit) representations. However, existing compression techniques typically require decompressing bounding boxes back to FP32 for intersection tests, thus failing to eliminate the computational burden and energy cost associated with complex FP32 arithmetic during BVH traversal. This work introduces AQB8, an RT accelerator designed to operate on a quantized BVH tree constructed using a novel multi-level quantization technique. This approach enables RT to operate directly on low-bit integers using simpler, area-efficient hardware units, thereby drastically reducing the need for FP32 arithmetic during BVH traversal while mitigating overheads associated with reduced precision. As a result, across various scenes, AQB8 achieves a 70% reduction in DRAM accesses, a 49% reduction in energy consumption, a 27% hardware area reduction, and a 1.82x performance speedup over modern GPU RT accelerators. Our code is publicly available at https://github.com/nycu-caslab/AQB8.