Careers

Teach the fleet to reason about production.

Thalamus SRE runs a governed fleet of agents against real incidents in real enterprise estates. The work below is about making those agents accurate enough to be trusted with production, and measurable enough to prove it.

Open roles

AI and ML engineering

All roles are remote-first. We care about demonstrated work more than credentials — if you have shipped models that other people depended on, tell us about that.

ML Engineer — SRE model fine-tuning

Remote · Full-time

Fine-tune open-weight models on incident data — traces, logs, change history, postmortems — so a specialist agent reaches a correct root cause more often than a general-purpose model does, on a fraction of the tokens.

  • Build and curate training sets from real incident telemetry
  • Fine-tune and quantise open-weight models for local and air-gapped deployment
  • Measure against held-out incidents, not vibes — regression on a benchmark blocks a release
  • Work within the token and latency budgets an on-call path actually allows

AI Engineer — custom LLM development

Remote · Full-time

Design the agents themselves: how an investigation is decomposed, what each specialist is allowed to conclude, and how their verdicts combine into an answer an engineer can act on at three in the morning.

  • Design multi-agent investigation workflows and their handoffs
  • Build retrieval over an estate's own runbooks, code and change records
  • Work across hosted and local models — customers in regulated sectors run air-gapped
  • Make agent decisions inspectable: every conclusion traceable to its evidence

ML Engineer — evaluation and guardrails

Remote · Full-time

Decide whether an agent is good enough to touch production. This is the role that says no, and it needs the evidence to make that stick.

  • Build evaluation harnesses over replayed real incidents
  • Measure accuracy, calibration and failure modes — including confident wrong answers
  • Design guardrails and approval gates that a regulated customer can audit
  • Track drift as models, estates and customer environments change

Interested? Write to careers@thalamusaicloud.com with something you have built and what you would want to work on here.

Investment

Thalamus AI Cloud is presently seeking aligned investors who are able to help expand the platform and its reach — enabling mid- and large enterprises to reduce toil and increase service reliability and effectiveness.

If that is work you want to back, use the form below, write to hello@madison-analytics.com, or call +1 (678) 430-1550.