Agent workflows
Building typed tool workflows that trace proposed state changes to bounded evidence and keep risky actions behind human review.
Building typed tool workflows that trace proposed state changes to bounded evidence and keep risky actions behind human review.
Building provenance, freshness checks, failure analysis, deterministic eval contracts, and privacy boundaries around model and retrieval behavior.
Building request harnesses, comparable artifacts, TTFT and throughput reports, KV-cache estimates, telemetry, and claim gates for serving experiments.
A career workspace that combines versioned resumes, recruiting history, and user corrections into evidence-backed recommendations, with human review before any update lands.
A Docker-sandboxed framework for reproducible tool-using agent evaluation with task manifests, JSONL trajectories, deterministic grading, replay, and failure analysis.
A canonical-first local recall layer using SQLite FTS5/BM25 over trusted files, with provenance, snippets, freshness checks, redaction, and workspace filtering.
An OpenAI-compatible serving benchmark toolkit for request workflows, TTFT, throughput, KV-cache estimates, quantization comparisons, and evidence-valid performance analysis.
For broader engineering work, review the full project list or the cloud and data systems page. For opportunities, contact Hanbin directly.