Run agents people can rely on.
Build, evaluate, secure, and operate agent systems with observable traces, cost controls, and clear failure boundaries.

- In planning
- No launch date is promised
- Remote-first
- Proposed delivery format
- Practical proof
- Capstone-led curriculum design
- Interest only
- Not an internship application
Why this belongs on the roadmap.
Agent demos are easy to produce and difficult to operate. The planned curriculum would make evaluation, observability, incident response, and security part of the build from the first prototype.
Research sources support the direction, not a launch date or outcome. This page does not promise a job, salary, client, income, or final curriculum.
Research basis
The work this curriculum would cover.
Each planned module is designed to produce evidence, moving from foundations to work a real operator can inspect.
- 01Choose between deterministic automation, one agent, and multi-agent designs
- 02Build tool-using agents with explicit permissions and schemas
- 03Improve prompts, retrieval, memory, and context using measured tests
- 04Create representative evaluation data from real traces
- 05Monitor latency, cost, errors, and task success
- 06Threat-model prompt injection, data leakage, and tool misuse
The proof this track would demand.
An agent system for a real business with an original tool integration, trace coverage, an evaluation suite in CI, cost and latency reporting, a threat model, alerts, and an incident runbook.

The working stack.
Tools can change before launch. The proposed workflow and quality bar are the durable part.
- Python
- Claude Agent SDK
- LangGraph
- MCP SDK
- Langfuse
- Braintrust
- GitHub
Register interest in this track.
Join the track-specific waitlist. This is separate from the internship intake waitlist and does not submit an application.