Projects · writing · sound

Stan Chen.

I make tools for remembering, practicing, and making sound.

Things that can be inspected.

03 projects

Tools behind the projects.

03 supporting pieces

A smaller set of work about verification, testing, and keeping agents useful under real constraints.

04

Evaluation & routing · Python

Inference-Time Scaling Engine

An OpenAI-compatible gateway that spends a bounded budget to generate, verify, repair, and select one response—with an inspectable trace.

Receipt: 342-case locked holdout · +4.97 points · 2.93× cost · scoped result, not a universal claim

Read the evidence

05

Quality & testing · Claude Code skill

Spec-Driven QA

A pipeline that turns a PRD into a classical test backbone, AI rubrics, executed checks, and a clear ship / hold decision.

Classical boundary coverage first; AI and OWASP probes only where they apply.

Read the skill

06

Reference · field manual

The AI Agent Handbook

Practical public guidance for building and operating AI agents, with clear receipts for what has been run, measured, or is still pending.

Build smaller · verify everything · keep the evidence visible

Open the handbook

Writing, in public.

Longer technical essays on Medium; looser notes only when they are ready to leave the notebook.

Dispatches / Substack

Infrequent notes, closer to the work.

Short updates, unfinished thoughts, and things worth passing on before they turn into an essay.

Visit Substack

Specific work beats a long list of claims. If something here is interesting, the link should let you see how it works.