01
AI infrastructure
Multi-GPU inference, request routing, prefix-cache behaviour and serving evaluation.
RadixGates · SGLang · llama.cpp · Docker
Penn Electrical Engineering · Queen’s Commerce + Computing
My work spans multi-GPU inference, local models and systems code, and now firmware and SoCs. I like getting a system running, then measuring the bottlenecks, failures and tradeoffs that abstractions hide.
Looking for Summer 2027 internships in AI infrastructure or embedded / hardware–software systems.
01
Multi-GPU inference, request routing, prefix-cache behaviour and serving evaluation.
RadixGates · SGLang · llama.cpp · Docker
02
C++ / Go systems work, concurrency, storage and performance measurement.
cdc-chunker · 7.4 GB/s on 8 threads
03
Bare-metal C, SoC profiling and digital IC fundamentals — a current focus at Penn.
ATmega328PB · Ultra96 · VLSI
04
PySpark / Hive pipelines, SQL, data quality and operational observability.
Shopee · 10+ jobs consolidated
Serving
Can a small Go gateway keep an SGLang pool serving when a node dies, and does its prefix routing still help on a real engine?
85.2% → 99.7% clean completions in a node-crash test, against the original gateway (4 simulated workers, 16 clients, 30 s)
Node-crash clean completions rose from 85.2% to 99.7% against the original, under the same injected faults; the test suite now has 55 tests.
On three real llama.cpp workers, the upgraded routing matched the original's cache hit rate (61.5% vs. 59.1%) rather than beating it, but cut the busiest worker's share of the prefill work from 73% to 43%.
Fine-tuning
Can a 0.5B model fine-tuned with LoRA answer questions about a database, on a laptop, without the data leaving it?
16.4% of answers ran and returned the wrong rows (Q4_K_M, 391 held-out questions, scored by executing the SQL)
Execution-correct answers rose from 42.7% to 82.6%; 16.4% ran cleanly but returned the wrong rows — invisible to a string-match check.
Guardrails only brought silent errors down to 9.1%, short of the 5% target, so the call was a human-reviewed suggestion tool, not an autonomous answerer.
Storage
How fast can content-defined chunking run in C++17 without changing where the chunks are?
7.4 GB/s on 8 threads, chunks identical to the sequential chunker (256 MiB in memory, 8 KiB average chunk, Apple M1)
About 2.0 GB/s on one thread, 7.4 GB/s on 8, output identical to sequential; 96.5% of a 67 MB source tree stayed reusable after 200 edits.
A merged LoRA model reused only 27.6% of its bytes against its base — ship the adapter, not the merged model.
Tooling
Given a model and some GPUs, how many do you need, why did the server not start, and are two benchmark runs even comparable?
6 factors differed between my RTX 4090 and A100 runs; the kit flags them instead of letting the comparison stand
Jun – Sep 2025
Big Data Engineering Intern
Shopee Logistics Network Tech Co.
Shenzhen, China
Mar – Jun 2025
Data Analyst Intern
Brix Technology Services
Canada
Jun – Sep 2024
Investment Analyst Intern
CITIC Securities
Beijing, China
Mar – Jun 2024
Investment Research Intern
Founder Securities
Beijing, China
Field notes
My gateway's failure tests all passed. A routing rule I had added still cut prefix-cache hits from 59.1% to 24.4% on real engines.
Read note →A fine-tuned 0.5B model writes the right SQL 82.6% of the time. What worries me is the 16.4% that runs cleanly and returns the wrong rows.
Read note →Only 27.6% of a fine-tuned model's bytes match its base, and the encoding mattered more than the chunker.
Read note →An edit is more likely to land in a big chunk, so the spread of chunk sizes matters as much as their average.
Read note →This semester at Penn: Smart Devices (bare-metal C on an ATmega328PB), SoC Architecture (profiling on an Ultra96) and Digital ICs and VLSI. More about me →
I’m looking for Summer 2027 roles in AI infrastructure or embedded / hardware–software systems.