Selected Engineering Work

AI/ML Systems and Software Engineering Projects

Evidence-backed work across agent workflows, model serving, evaluation, recommendation systems, retrieval, backend platforms, and cloud infrastructure.

Loading projects...

6
Featured Projects
6
Other Projects
2
Demos and Products

Filter by Technology

Featured Projects

Case-study format: problem, build, stack, and impact.

CareerOS

Private product with public evidence

Repo: public-careeros

Next.jsTypeScriptGemmaAI Agents

Problem

Recruiting updates arrive through email, while trackers and follow-up plans drift without bounded evidence and review.

What I Built

Built Gmail triage, typed workflow extraction, confidence-aware review, deterministic fallback, and model traces over 130 versioned sanitized fixtures.

Impact

The deterministic product-loop regression passes 118 of 130 contracts (90.8%) with 0 unsafe automatic mutations; 12 false review routes remain documented in checked-in error analysis.

LLM Inference Acceleration Playground

Public benchmark toolkit

Repo: llm-inference-acceleration-playground

PythonLLM ServingEvaluationGPU Telemetry

Problem

Serving comparisons become misleading when request modes, artifacts, hardware evidence, or run completeness differ.

What I Built

Built streaming and non-streaming clients, concurrent open-loop and closed-loop workflows, JSONL/CSV artifacts, manifests, latency reports, KV-cache estimates, quantization comparisons, optional GPU telemetry, and claim-audit gates.

Impact

The toolkit rejects mock, incomplete, or incomparable runs from performance rankings and preserves the evidence needed to audit each comparison.

Agent Sandbox Eval

Public evaluation framework

Repo: agent-sandbox-eval

PythonDockerAI AgentsEvaluation

Problem

Tool-using agent experiments need reproducible tasks, isolated execution, and controls that expose harness failures.

What I Built

Built 25 terminal, SWE-style, and stateful-tool tasks with deterministic grading, replay, read-only container roots, dropped Linux capabilities, seccomp, no-new-privileges, PID limits, and explicit writable mounts.

Impact

The scripted oracle and deterministic local control each score 25/25; the no-op negative control scores 0/25.

Boring Agent Memory

Public memory infrastructure

Repo: boring-agent-memory

PythonSQLiteFTS5Retrieval

Problem

Agents need fast local recall without replacing canonical files or leaking secrets across workspaces.

What I Built

Built SQLite FTS5/BM25 retrieval with provenance, source snippets, freshness verification, secret redaction, workspace filtering, and CLI, Python, and stdio interfaces.

Impact

On 120 sanitized synthetic queries, the benchmark records 0.990 Recall@1, 0.995 MRR, 1.000 no-answer precision, and 0 privacy leaks; exact-phrase grep reaches 0.200 Recall@1.

Sequential Recommendation Benchmark

Team offline benchmark
PythonPyTorchRecommendation SystemsSASRec

Problem

Sequential recommendation models need a shared split and metrics before their ranking quality can be compared.

What I Built

Contributed to preprocessing, model implementation, training, evaluation, comparison, and reporting across 23,936 users, 2.62M interactions, and 50K items with leave-one-out evaluation at K=50.

Impact

SASRec reaches HR@50 0.1525 and NDCG@50 0.0503, about 4.8x and 5.4x the MostPopular baselines.

cc-lite

Public research baseline

Repo: cc-lite

PythonPyTorchMCTSXiangqi

Problem

Xiangqi research needs a correct rules engine and reproducible training and evaluation workflows before model strength can be studied.

What I Built

Built an 8,100-action move space, legal policy masking, a PyTorch policy-value CNN, PUCT MCTS, self-play, checkpoint training, engine distillation, and CPU, MPS, and CUDA evaluation workflows.

Impact

The repository provides a reproducible baseline for rules, search, training, and engine comparison work.

Other Projects and Experiments

Smaller builds, coursework, prototypes, and engineering practice.

SFML Graphing Calculator

Repo: Graphing_Calculator_SFML

A C++17 and SFML graphing calculator with expression parsing, RPN evaluation, interactive plotting, expression history, and headless tests.

C++SFMLParsingComputer Graphics

Tripemini

A 48-hour multimodal food and travel MVP with bounded Gemini inference, structured TypeScript outputs, validation, deadlines, rate limits, and concurrency controls.

TypeScriptNext.jsGeminiVercel

OAuth 2.0 Security Reference

Repo: Next.js-OAuth2-Template

A Next.js authentication reference covering verified provider identity, PKCE and state, callback replay, session rotation, rate limits, and email-collision regression tests.

TypeScriptNext.jsOAuth2Security

Website Image Downloader

Repo: Website-Image-Downloader

A bounded Next.js ingestion service with DNS pinning, redirect revalidation, SSRF and DNS-rebinding defenses, private-address blocking, timeouts, and deterministic security tests.

TypeScriptNext.jsSecurityAutomation

SQL in C++

Repo: SQL

A C++17 SQL engine with 4 KB checksummed pages, persistent per-field B+ tree indexes, WAL-backed atomic writes, copy-first migration, and crash-recovery tests.

C++SQLStorage EngineDatabase Systems

Mail Guard

Repo: mail-guard

A team IoT mailbox-monitoring prototype with ESP32-CAM firmware, a Next.js backend, image storage, and alert workflows.

Next.jsTypeScriptC++IoT

Interested in collaborating?

I am open to software engineering, AI/ML systems, and applied AI opportunities.

Get in Touch →