Yuchen He 中文

Penn Electrical Engineering · Queen’s Commerce + Computing

I build systems behind AI — and I’m moving closer to the hardware.

My work spans multi-GPU inference, local models and systems code, and now firmware and SoCs. I like getting a system running, then measuring the bottlenecks, failures and tradeoffs that abstractions hide.

Looking for Summer 2027 internships in AI infrastructure or embedded / hardware–software systems.

See the code on GitHub Email

A rule that looked harmless

Originalhash % N, queues when full59.1%
Spill when fullthe upgrade's first rule24.4%
Placement + waitthe fix61.5%
Round-robinno gateway22.0%
Three llama.cpp workers on an Apple M1, 120 requests over six shared prompts, cache limited to the slots, median of 3 runs.

Capabilities

01

AI infrastructure

Multi-GPU inference, request routing, prefix-cache behaviour and serving evaluation.

RadixGates · SGLang · llama.cpp · Docker

02

Systems & performance

C++ / Go systems work, concurrency, storage and performance measurement.

cdc-chunker · 7.4 GB/s on 8 threads

03

Embedded & HW–SW

Bare-metal C, SoC profiling and digital IC fundamentals — a current focus at Penn.

ATmega328PB · Ultra96 · VLSI

04

Data engineering

PySpark / Hive pipelines, SQL, data quality and operational observability.

Shopee · 10+ jobs consolidated

Selected work

  1. Fine-tuning

    llm-finetune-lab

    Can a 0.5B model fine-tuned with LoRA answer questions about a database, on a laptop, without the data leaving it?

    16.4% of answers ran and returned the wrong rows (Q4_K_M, 391 held-out questions, scored by executing the SQL)

    PyTorchPEFTDDPllama.cppSQLite

    See the measurement
    Measured

    Execution-correct answers rose from 42.7% to 82.6%; 16.4% ran cleanly but returned the wrong rows — invisible to a string-match check.

    Judgment

    Guardrails only brought silent errors down to 9.1%, short of the 5% target, so the call was a human-reviewed suggestion tool, not an autonomous answerer.

  2. Storage

    cdc-chunker

    How fast can content-defined chunking run in C++17 without changing where the chunks are?

    7.4 GB/s on 8 threads, chunks identical to the sequential chunker (256 MiB in memory, 8 KiB average chunk, Apple M1)

    C++17CMakeGitHub Actions

    See the measurement
    Measured

    About 2.0 GB/s on one thread, 7.4 GB/s on 8, output identical to sequential; 96.5% of a 67 MB source tree stayed reusable after 200 edits.

    Judgment

    A merged LoRA model reused only 27.6% of its bytes against its base — ship the adapter, not the merged model.

  3. Tooling

    llm-serving-eval-kit

    Given a model and some GPUs, how many do you need, why did the server not start, and are two benchmark runs even comparable?

    6 factors differed between my RTX 4090 and A100 runs; the kit flags them instead of letting the comparison stand

    Python

Experience

Jun – Sep 2025

Big Data Engineering Intern

Shopee Logistics Network Tech Co.

Shenzhen, China

  • Consolidated 10+ SQL jobs into a dependency-aware PySpark/Spark SQL batch layer on Hive, standardizing incremental and snapshot patterns and cutting end-to-end runtime by about 40%.
  • Unified algorithm and regional forecasting schemas into a compute-once, multi-consume model; partitioning and predicate pushdown cut scanned data per query by about 60%.
  • Built WMAPE model-version dashboards and automated checks for null rate, volume drift and freshness; investigated pipeline and data-quality anomalies with algorithm and operations teams.

Mar – Jun 2025

Data Analyst Intern

Brix Technology Services

Canada

  • Built an Excel VBA and Power Query analytics workflow for 10,000+ case records — rule-based deduplication, open/closed/reopened status classification and automated aggregation — improving reporting speed about 25%.
  • Analyzed three years of sales data from 1,000+ retail stores in MySQL, surfacing seasonal patterns and margin peaks, and built Power BI dashboards broken down by store, vendor and manager.

Jun – Sep 2024

Investment Analyst Intern

CITIC Securities

Beijing, China

  • Supported IPO due diligence, valuation materials and disclosure cross-checks — tracing a conclusion back to its evidence, a habit that now carries into how I write up benchmarks.

Mar – Jun 2024

Investment Research Intern

Founder Securities

Beijing, China

  • Turned recurring market, industry, and ESG research into refreshable Power Query and VBA workflows; building the tool turned out to be as interesting as the analysis itself.
  • Modeled multi-source financial and market data in Excel and Wind for KPI benchmarking and trend analysis, including multi-period ROE, margin and leverage analysis.

Field notes

Notes

  1. 01

    A simulation can't see a cache

    My gateway's failure tests all passed. A routing rule I had added still cut prefix-cache hits from 59.1% to 24.4% on real engines.

    Read note →
  2. 02

    The wrong answers that raise no error

    A fine-tuned 0.5B model writes the right SQL 82.6% of the time. What worries me is the 16.4% that runs cleanly and returns the wrong rows.

    Read note →
  3. 03

    Same weights, different dtype

    Only 27.6% of a fine-tuned model's bytes match its base, and the encoding mattered more than the chunker.

    Read note →
  4. 04

    Average chunk size isn't what an edit costs

    An edit is more likely to land in a big chunk, so the spread of chunk sizes matters as much as their average.

    Read note →

Now

This semester at Penn: Smart Devices (bare-metal C on an ATmega328PB), SoC Architecture (profiling on an Ultra96) and Digital ICs and VLSI. More about me →

Let’s talk about systems.

I’m looking for Summer 2027 roles in AI infrastructure or embedded / hardware–software systems.