Projects

Listed in reverse chronological order.

2026

  1. WG-KV: Learned KV Admission for Efficient Long-Context LLM Inference
    WG-KV: Learned KV Admission for Efficient Long-Context LLM Inference
    Built the training and serving stack behind our EMNLP 2026 paper, which introduces KV Admission alongside KV Selection and Eviction. Trained a lightweight per-head write gate by self-distillation from the frozen full-attention model, then built a dual-cache paged manager and sparse attention kernels over the resulting ragged caches, turning theoretical sparsity into 2.6–4.2x prefill and 1.6–2.6x decode speedups.
    Master's Thesis, NTU, 2026
  2. STM32 Deployment and Benchmarking Toolchain for Multi-Branch TinyML
    STM32 Deployment and Benchmarking Toolchain for Multi-Branch TinyML
    Built the deployment and benchmarking toolchain for our TCAD 2026 paper on DupNAS, a framework that jointly optimizes neural architecture and multi-branch tensor splitting to fit vision models onto resource-constrained microcontrollers. The toolchain turns the searched networks into real STM32 measurements via tile-by-tile ONNX rewriting and deployment through TensorFlow Lite Micro and STM32Cube.AI.
    Research Project, Embedded and Mobile Computing Lab (Prof. Pi-Cheng Hsiu), Academia Sinica, 2026

2025

  1. Linux Kernel Programming: Syscall Interception and Page-Table Manipulation
    Linux Kernel Programming: Syscall Interception and Page-Table Manipulation
    Built a loadable kernel module that hides itself from the module list, masquerades process names, and filters syscalls per-process by hooking the syscall table. Also added a custom ARM64 Linux syscall that exposes a process's page tables to user space, and built a userfaultfd handler that walks the page-table hierarchy to translate addresses and install page-table entries on demand.
    Course Team Project, Advanced Topics in Software Systems Design and Implementation (Prof. Shih-Wei Li), NTU, 2025
  2. AQB8 Ray Tracing Accelerator: From Algorithm to FPGA and ASIC
    AQB8 Ray Tracing Accelerator: From Algorithm to FPGA and ASIC
    Led the algorithm and hardware design behind our ISCA 2025 paper on AQB8, a ray tracing (RT) accelerator that replaces FP32 with low-bit integer arithmetic during BVH traversal. Prototyped it on a Xilinx Zynq FPGA with Vitis HLS, Vivado, and PetaLinux, then took the same design through a TSMC 40nm standard-cell ASIC flow with Catapult HLS, Design Compiler, QuestaSim, and PrimePower, for 49% lower energy and 27% less area than modern GPU RT accelerators.
    Research Project, Computer Architecture System Laboratory (Prof. Tsung Tai Yeh), NYCU, 2025

2024

  1. Multimodal Perception and Comprehension of Corner Cases in Autonomous Driving
    Multimodal Perception and Comprehension of Corner Cases in Autonomous Driving
    Led the team to 1st place (out of 12 teams) on multimodal understanding of autonomous driving corner cases. Enhanced LLaVA with LoRA fine-tuning plus cross-attention layers that fuse segmentation, instance, depth, and ViT features, enabling the model to reason over complex driving scenarios. Also built a self-critiquing augmentation loop where a frozen LLaVA compares the fine-tuned model's predictions with the ground truth and turns each missed detail into a new QA pair targeting the model's blind spots.
    Course Team Project, Deep Learning for Computer Vision (Prof. Yu-Chiang Frank Wang), NTU, 2024
  2. Dataflow Graph Extraction in LLVM for High-Level Synthesis
    Dataflow Graph Extraction in LLVM for High-Level Synthesis
    Built a custom LLVM compiler pass that turns C++ functions into dataflow graphs, making explicit the computational dependencies that hardware designers rely on to pipeline a design. Working on LLVM's intermediate representation rather than the syntax tree lets the pass compose with existing optimizations such as dead code elimination (DCE) and common subexpression elimination (CSE).
    Course Project, Advanced Compiler Design (Prof. Shih-Wei Liao), NTU, 2024
  3. Design and Implementation of a RISC CPU: From RTL to GDSII
    Design and Implementation of a RISC CPU: From RTL to GDSII
    Independently designed a 16-bit single-core CPU with a MIPS-like ISA and AXI4 interfaces for instruction and data memory, then carried it through the entire ASIC design flow from RTL to GDSII, covering logic synthesis, floorplanning, automatic place & route (APR), and final layout generation. Verified with RTL, gate-level, and post-layout simulation, plus STA, LVS, and DRC using industry-standard CAD tools.
    Course Project, Integrated Circuit Design Laboratory (Prof. Chen-Yi Lee), NYCU, 2024
    Source code is not publicly available due to course policy.

2023

  1. Domain-Specific Accelerator for MLP Inference on an FPGA SoC
    Domain-Specific Accelerator for MLP Inference on an FPGA SoC
    Designed a domain-specific accelerator that offloads handwritten-digit MLP inference from an FPU-less soft-core CPU on a Xilinx Artix-7 FPGA. Reordered the computation to keep the floating-point multiply-adder saturated, replicated the datapath for data parallelism, and used BRAM-backed FIFOs to overlap host transfers with computation, resulting in a combined 309× speedup over the software-only baseline.
    Course Project, Microprocessor Systems: Principles and Implementation (Prof. Chun-Jen Tsai), NYCU, 2023
  2. Parallel Ray Tracer with SIMD, OpenMP, MPI, and CUDA
    Parallel Ray Tracer with SIMD, OpenMP, MPI, and CUDA
    Built a ray tracer from scratch, then parallelized it with SIMD intrinsics, OpenMP, MPI, and CUDA, guided by a loop-level analysis of where parallelism actually pays off across pixels, samples, bounces, objects, and vector components. Tuned the CUDA implementation further with coalesced memory access and a split-kernel design that lowers register pressure to improve occupancy.
    Course Team Project, Parallel Programming (Prof. Yi-Ping You), NYCU, 2023
  3. Memory-Efficient Attention Operator for TensorFlow Lite Micro
    Memory-Efficient Attention Operator for TensorFlow Lite Micro
    Investigated the challenges of deploying Transformer models on microcontrollers (TinyML) through TensorFlow Lite for Microcontrollers (TFLM), and identified the Multi-Head Attention layer as the critical SRAM bottleneck during on-device inference. Developed and implemented a custom TFLM operator that reduces peak SRAM requirement by over 90% (from 17MB to ~1MB) for the MobileViT-XXS model.
    Summer Internship, Embedded and Mobile Computing Lab (Prof. Pi-Cheng Hsiu), Academia Sinica, 2023
  4. Mixed-Precision Ray Tracing for Mobile Platforms
    Mixed-Precision Ray Tracing for Mobile Platforms
    Collaboration with MediaTek on reducing the memory bandwidth demand of ray tracing (RT) accelerators on mobile platforms. Built a ray tracer supporting physically-based rendering and parallel BVH traversal, then designed a novel mixed-precision traversal algorithm validated via SystemC TLM simulation, achieving 20% memory savings for key RT data structures.
    Research Project, NYCU and MediaTek Advanced Research Center (MARC), 2023
    Source code is not publicly available due to NDA restrictions.
  5. Channel, Global, and Detailed Routers for VLSI
    Channel, Global, and Detailed Routers for VLSI
    Implemented three routers spanning the VLSI routing flow: a greedy channel router, a global router driven by negotiation-based rip-up and reroute, and a detailed router that turns routability into a CNF formula solved by a MaxSAT solver. Together they trace how routing scales, from complete search on small grids to heuristics that handle the ISPD 2008 contest benchmarks.
    Course Project, The Applications of Algorithms on Routing Problems (Prof. Yih-Lang Li), NYCU, 2023
  6. Linux API Sandbox and Instruction-Level Debugger
    Linux API Sandbox and Instruction-Level Debugger
    Built two low-level Linux tools from scratch. The first enforces an API blacklist on unmodified binaries, preloading a shared library that hijacks __libc_start_main() and parses the ELF to rewrite GOT entries, so file, network, and shell calls can be interposed. The second is a ptrace-based debugger with on-the-fly disassembly, gdb-style int3 breakpoints, and a time-travel command backed by fork() snapshots.
    Course Project, Advanced Programming in the UNIX Environment (Prof. Chun-Ying Huang), NYCU, 2023

2022

  1. Mini-Pascal Compiler: From Source Code to JVM Bytecode
    Mini-Pascal Compiler: From Source Code to JVM Bytecode
    Independently designed a compiler for Mini-Pascal, a Pascal subset, covering a flex-based scanner, a bison LALR parser that builds the AST, semantic analysis over a scoped symbol table with static type checking, and a code generator emitting Jasmin assembly that is further assembled into JVM bytecode. The resulting compiler correctly compiles and runs non-trivial programs like quicksort and recursive Fibonacci.
    Course Project, Introduction to Compiler Design (Prof. Wuu Yang), NYCU, 2022
  2. Real-time Parking Vacancy Forecasting System
    Real-time Parking Vacancy Forecasting System
    Led the team to Best Project (out of 249 students) with a system that forecasts vacancy at New Taipei City's public parking lots in real time. Designed the full architecture covering automated data collection, time-series model training, and a Vue frontend with interactive forecast charts. The LSTM model achieved low prediction error, with serverless inference on Azure Functions fast enough for live requests.
    Course Team Project, Introduction to Artificial Intelligence (Prof. Wen-Chih Peng), NYCU, 2022

2021

  1. Object Detection of Water-Holding Containers for Dengue Mosquito Control
    Object Detection of Water-Holding Containers for Dengue Mosquito Control
    Ranked 1st (out of 27 teams) on an AIdea object detection challenge for locating water-holding containers, the breeding sites of dengue-carrying mosquitoes. Fine-tuned COCO-pretrained YOLOv5 detectors and synthesized extra training images by pasting cut-out objects onto backgrounds, countering the severe class imbalance in the provided data.
    Course Project, Introduction to Deep Learning and Applications (Prof. Hung-Hsun Chen), NYCU, 2021

2019

  1. Web-Based Search Tool for College Admission Requirements
    Web-Based Search Tool for College Admission Requirements
    Built a web tool that lets Taiwanese high school students look up which entrance-exam subjects and score thresholds each university department counts toward admission, consolidating information that is otherwise scattered across hundreds of pages on the official admission committees' sites. Since 2019, it has served 64,000 page views from 24,000 unique visitors.
    Side Project, 2019