Own ML projects end-to-end: take problems from exploration through production deployment and keep iterating once real users are on them
Build agentic LLM systems: design multi-step workflows with tool use, retrieval, and orchestration, and make them reliable enough for clinical settings
Treat evaluation as core engineering work: build eval sets, offline and online harnesses, LLM-as-judge pipelines with human review, and regression tests that catch quality drops before they reach users
Improve model quality with whatever fits the problem: prompting, retrieval, distillation, or fine-tuning, chosen on evidence rather than habit
Work across the full AI stack: data prep, model adaptation, serving, monitoring, and the feedback loops that keep systems improving in production
Partner with Product, Clinical, and Engineering: translate clinical requirements into technical decisions and surface tradeoffs early
Help the team get better: review code, share what you learn, and mentor engineers earlier in their careers
Benefits
A stimulating, fast-paced environment with lots of room for creativity
A bright future at a promising high-tech startup company
The opportunity to work with a talented team and to add real value to an innovative solution with the potential to change the future of healthcare
A stimulating environment with room for creativity
A flexible environment where you can control your hours (remotely) with unlimited vacation
Access to our health and well-being program (digital therapist sessions)
Remote or Hybrid work policy
Career development and growth- Hands-on LLM work in production: prompting, retrieval, tool calling, and agent-style workflows
Clear communication with both technical and clinical stakeholders
Experience shipping ML systems to production that people actually depend on
Comfort with ambiguity: you’ve taken loosely defined problems and turned them into something running in production
Strong ML fundamentals: you know which approach fits which problem and can reason clearly about tradeoffs
A rigorous approach to evaluation: you’ve built eval datasets and frameworks, and you can tell a real improvement from noise
Solid engineering skills: production-quality code, familiarity with distributed systems, and the patience to debug messy ML pipelines
Experience with fine-tuning or preference optimization (RLHF, DPO, or similar)
Healthcare AI, or other high-stakes domains where errors carry real cost
Built agent frameworks or evaluation tooling from scratch
Open source contributions, technical writing, or other knowledge sharing