logo
Yubo Cao

CTO & Co-founder · Similate · a16z Speedrun SR7

Yubo Cao

I build swarms of LLM agents at Similate — and research how to evaluate them and keep them safe.

Caltech CS '29 · Student Researcher, Anima AI + Science Lab · prev. Lockheed Martin AI/ML

1000s
concurrent LLM agents per simulation
10+
studios running on Similate
NeurIPS '26
first-author paper under review
~10⁶×
faster inference than Monte Carlo

Now

What I'm building right now

Similate — CTO & Co-founder

Thousands of synthetic users, one simulation

Similate (a16z Speedrun SR7) simulates real user populations with swarms of LLM agents to automate A/B testing — thousands of agents running concurrently per simulation, in production for 10+ design and brand-naming studios.

Multi-agent systemsSynthetic usersa16z Speedrun SR7

Research — LLM Safety & Evaluation

How do agent swarms fail?

Running thousands of production LLM agents is a front-row seat to how they break. My current research asks how to evaluate and align agent systems: benchmarks for synthetic-user fidelity, and fine-tuning studies where training loss improves while true ranking ability collapses — evidence that you must evaluate the capability you care about, not its proxy.

AI safetyAgent evalsFine-tuning diagnostics

First-author paper — NeurIPS 2026 (under review)

Learning to outrun Monte Carlo

PTNO trains neural-operator surrogates for particle transport on cheap, noisy Monte Carlo labels — and ends up more accurate than its own training data, with inference up to ~10⁶× faster than converged Monte Carlo. Trained on 1M+ simulated configurations.

Neural operators~10⁶× inference speedup1M+ configurations

Side project — Live trading stack

Sub-second from social signal to order

Jiedusuo is a signal-relay and auto-trading platform I built and operate: an Axum Rust server on Postgres, a WebSocket feed, and a headless Python execution daemon, serving thousands of users at ~200k requests/day with ~0.5 s p50 end-to-end entry latency. The backtest lab behind it pre-registers every experiment, applies hard gates before a single maximand, and scores per-author edge with a recency-weighted EWMA with shrinkage before any strategy goes live.

Rust · Axum · Postgres~200k requests/day~0.5 s p50 latency

Education

Caltech CS
Caltech

Caltech

Class of 2029, Computer Science. GPA 4.2/4.0. Learning Systems (A+), Decidability & Tractability (A), Experimental Robotics (A).

Georgia Tech
Georgia Tech

Georgia Institute of Technology

Dual Enrolled. GPA: 4.0. CS 1331 (A): Intro to Object-Oriented Programming; MA 1554 (A): Linear Algebra; MA 2551 (A): Multivariable Calculus; MA 3012 (A): Applied Combinatorics; and MA 2552 (A): Differential Equations.

GSMST
GSMST

Gwinnett School of Math, Sci, and Tech

GPA: 4.4, NGA: 101.960, Salutatorian. AP Computer Science A (5/5), AP Statistics (5/5), AP Computer Science Principles (5/5), AP Calculus BC (5/5), AP Physics C and E&M (5/5)

Experience

CTO & Co-founder
Similate · a16z Speedrun SR7 · Apr 2026—now
  • Building a multi-agent synthetic-user simulation platform for automated A/B testing, running thousands of LLM agents concurrently per simulation.
  • In production for 10+ clients across design and brand-naming studios; evaluation benchmarks curated from SimBench, with controlled experiments run before any customer-facing metric ships.
Anima AI + Science Lab, Caltech · advisor Prof. Anima Anandkumar logo
Student Researcher
Anima AI + Science Lab, Caltech · advisor Prof. Anima Anandkumar · Oct 2025—now
  • First-author work (PTNO, under review at NeurIPS 2026): trained neural-operator surrogates for Monte Carlo particle transport from noisy, low-sample MC labels across 1M+ configurations, with inference up to ~10⁶× faster than converged Monte Carlo.
  • Diagnosed the Jensen bias that appears when noisy labels are pushed through a log transform, and designed a positivity-preserving operator head and loss that predict fields spanning 9–11 orders of magnitude from noisy labels alone.
  • Built the data-generation and training pipeline (Slurm, CUDA/NVIDIA Warp) behind the paper's 1M+ simulated configurations.
Lockheed Martin · Information Security Office, then Guidance, Navigation & Control logo
Cybersecurity & GNC Intern
Lockheed Martin · Information Security Office, then Guidance, Navigation & Control · May—Dec 2025
  • Engineered "Security Advisor," a GenerativeAI assistant that won 1st Place in enterprise-wide Intern AI challenge. Projected to save ~28,000 engineering hours (approximately $2M+/yr) annually by consolidating security protocols into a unified, accessible portal.
  • Built GenGNC, a LangChain multi-agent system that turns a high-level concept of operations into requirements and Simulink simulation prototypes, projected to cut prototyping time by 90%.
  • Built emailing system for IT compliance tasks, Q&A system for controlled unclassified information handling, and AI agent for compliance audit, saving 80+ hrs/person/yr (approximately 400+ hrs/yr).
  • Created application to fetch proposals to change regulations (dockets), saving 32+ hrs/docket (approximately 768+ hrs/yr).
Lockheed Martin, Chief of Data Analytics Office logo
AI/ML Engineer Intern
Lockheed Martin, Chief of Data Analytics Office · May 2023—Aug 2024
  • Migrated legacy Domino system to AI Factory, saving $2M+ in subscription costs and introducing MLOps practices in alignment with 1LMX.
  • Rearchitected corporate-wide website from Angular to WordPress, saving 120 hrs/yr in maintenance costs.
  • Fine-tuned BERT & used bipartite matching to match candidates to roles in internal leadership development program, saving 56 hrs/yr.
  • Developed retrieval-augmented discrepancy reports system using StreamLit and AI factory for Solumina, automating discrepancy reports.
GSoC
GSoC

Open Source Contributor

Google Summer of Code (GSoC) · Jun 2025—now

  • Selected as a contributor (top 1,300 globally) for JabRef (4k+ stars) with a $3,000 stipend.
  • Enhanced onboarding experience, built declarative API for UI creation (similar to Tipkits), and event-driven rendering system in JavaFX.
Medlytics
MIT BWSI

Medlytics Student

MIT Beaver Works Summer Institute · Jun—Aug 2024

  • Developed medical image visual question and answering system through CLIP and Word2Vec (ImageCLEF2023), achieving 80%+ accuracy and surpass the State-of-the-Art accuracy by 5%.
HackGwinnett
HackGwinnett

Founder & President

HackGwinnett · Oct 2022—Oct 2024

  • Organized largest high school hackathons in Georgia three times ('23, '24, '25), with 200+ attendees, 50+ schools, and 10+ counties each.
  • Hosted CS days for K-5 (200+ attendees) and raised $2,000+ partnering with State Farm & Amazon.
GSMST
GSMST

Independent Researcher

Gwinnett School of Math, Sci, and Tech · Sep 2022—May 2025

Skills

Languages

PythonCC++RustJavaTypeScriptJavaScriptSQLKotlin

AI/ML

PyTorchLangChainHugging FaceLoRA fine-tuningscikit-learnPandasRAG

Numerical & HPC

NumPySciPyMonte Carlo methodsneural operatorsCUDANVIDIA WarpSlurm

Web

Next.jsReactReact NativeExpoTailwindAxumFlaskSQLAlchemy

Cybersecurity

LinuxBashsystem hardeningIDACTFs

Infra & Tools

AWS (Bedrock)VercelSupabasePostgresDockerKubernetesTemporalGitHub ActionsGit

Awards & Honors

check_circle

CalHacks 12.0 (UC Berkeley): Most Creative Hack, 1st of 695 teams ('25)

check_circle

Carnegie Mellon University picoCTF: Top 8 of 1,498 US teams ('23); top 5% in US ('24)

check_circle

LM CodeQuest: Top 3 in nation ('23) & 1st in state ('25); LM CyberQuest: 1st in state ('24, '25)

check_circle

Air & Space Forces Association CyberPatriot: 4th in nation & 1st in state ('23); top 1% in nation ('24)

check_circle

Riley's Way Call For Kindness: Fellow (top 50 globally), $5K grant for JobNest, a job platform for immigrants

check_circle

Georgia Governor's Honors Program: Finalist for Software Engineering (top 13 statewide)

check_circle

Coolidge Senator: Top 100 out of over 4,200 global applicants with $1,000 Scholarship ('24)

check_circle

USACO: Silver division