SRE-as-a-Service

Your own SRE fleet, running by Friday.

A governed fleet of AI agents that investigates, remediates and documents incidents across the enterprise stack — cutting mean-time-to-resolution from 4 hours to 47 minutes, in your own instance, in the region you choose. Provisioned in minutes, not a quarter of procurement.

Create a low-cost Staging instance to start. Monthly or annual after that — no long procurement either way.

70%
ITSM toil removed
47min
MTTR achieved
100×
Services per team
85%
Tickets automated
What it does
Autonomy

Agents monitor golden signals, investigate incidents and remediate on production code — alert to fix in minutes.

Governance

RBAC, approval gates, human escalation and immutable provenance for every agent decision.

Security

Runs on major US closed models or a local LLM; secrets stay server-side and your data never leaves your instance.

How it works

Four steps, and one of them is waiting.

The deployment packs are already built and tested against live accounts in all three major clouds. Provisioning runs them; it does not invent anything on the day.

Create an account

Email, password, confirm the address. No sales call to get in.

about a minute

Choose a name and a region

Your instance lives at yourname.sre-prod.thalamusaicloud.com, in the region closest to the systems you are monitoring.

under a minute

We provision it

Infrastructure, database, hardened defaults, DNS. You watch the actual steps — not a spinner — and we email you when it is ready.

a few minutes

Connect your first service

Point it at your monitoring and your repositories. The fleet starts watching what you onboard, and nothing else.

the same afternoon
What you get

An agent fleet, not a dashboard.

Dashboards tell you something broke. Thalamus works out why, writes it down, and proposes the change — with the evidence attached.

01

Investigation

Correlates signals across monitoring, logs, traces and recent changes, then says what it thinks caused the incident and how confident it is.

02

Resolution

Proposes a fix as a pull request against your repository, with the test that demonstrates it. You review and merge — or do not.

03

Evidence

Every claim is labelled as fact, observation, inference or hypothesis. You can tell what was measured from what was guessed.

04

Data protection

Redaction, retention and subject-rights handling are built in, not a later module. Credentials are encrypted at rest with your own key.

05

Guarded autonomy

Observe, recommend, approve — then more if you want it. The fleet does not act beyond the level you set, and the level is yours to move.

06

Your cloud, your region

AWS, Azure or GCP. Data residency is a dropdown at signup, not a professional services engagement.

Pricing

Priced by the services you actually onboard.

Not per seat, not per host, not per gigabyte ingested. A service is a production thing you want the fleet to watch and act on — its own deployment, its own golden signals, its own on-call expectation. Not every repository, and not every container.

Thalamus SRE

Priced per onboarded service
Staging
Try it · up to 2 services
$250
for one week
Start a test
  • Full agent fleet
  • Your own instance and database
  • Promote and keep your setup

Then it pauses — promote or stop. Nothing is deleted.

Launch
Up to 3 services
$999
per month
or $10,788/year — save 10%
  • Full agent fleet
  • Shared infrastructure
  • Onboarding included

If needed, Thalamus Advisory can help tune your service landscape.

Get started
Team
4 to 10 services
$2,998
per month
or $32,376/year — save 10%
  • Everything in Launch
  • Priority support
  • Change-window pinning

If needed, Thalamus Advisory can help tune your service landscape.

Get started
Division
11 to 18 services
$5,998
per month
or $64,776/year — save 10%
  • Everything in Team
  • Multiple instances
  • SSO and audit export

Includes a three-month Thalamus Advisory engagement to tune your service landscape.

Get started
Enterprise
19+ services
$9,998
per month
or $107,976/year — save 10%
  • Dedicated AWS account
  • Own database
  • Any region

Includes a three-month Thalamus Advisory engagement to tune your service landscape.

Talk to us

ThalamusTrace

Priced per onboarded service, on the same tiers
Staging
Try it · up to 2 services
$50
for one week
Start a test
  • Metrics, traces, logs and problems
  • Your own instance and database
  • Promote and keep your setup

Then it pauses — promote or stop. Nothing is deleted.

Launch
Up to 3 services
$199.80
per month
or $2,157.84/year — save 10%
  • OTLP-native ingest
  • Shared infrastructure
  • Onboarding included

If needed, Thalamus Advisory can help tune your service landscape.

Get started
Team
4 to 10 services
$599.80
per month
or $6,477.84/year — save 10%
  • Everything in Launch
  • Priority support
  • Alert receivers and webhooks

If needed, Thalamus Advisory can help tune your service landscape.

Get started
Division
11 to 18 services
$1,199.80
per month
or $12,957.84/year — save 10%
  • Everything in Team
  • Multiple instances
  • SSO and audit export

Includes a three-month Thalamus Advisory engagement to tune your service landscape.

Get started
Enterprise
19+ services
$1,499.80
per month
or $16,197.84/year — save 10%
  • Dedicated AWS account
  • Own database
  • Any region

Includes a three-month Thalamus Advisory engagement to tune your service landscape.

Talk to us

Each product is bought on its own tier, so an estate can run one, the other, or both. Running both puts the platform that raises a problem and the fleet that investigates it under one boundary — see Solutions for how they connect.

What this replaces

The usual alternative is three or four products and someone to hold them together: an APM priced per host, an ITSM priced per seat, a log platform priced per gigabyte, and a standing consulting engagement to keep the integrations from drifting apart. Estates that size routinely spend millions a year on that arrangement — and still wake a human at 3am to read the dashboards it produces.

Thalamus is one vendor, one boundary, and a price published on this page. Both products at the largest tier come to $137,973.60 a year, or $124,173.84 billed annually — for a fleet that does the reading, the correlating and the first draft of the fix.

No integration project, no per-host arithmetic, and nothing bandaged together after the fact.

Division and Enterprise: three months of Thalamus Advisory

Both tiers include a limited three-month engagement with Thalamus Advisory to tune your service landscape for architecture and performance — which services genuinely warrant their own SLOs, where the golden signals should sit, and which dependencies are load-bearing enough to page on.

It is deliberately time-boxed. The point is to leave your team able to onboard the next service without us, not to become the standing consulting engagement this page argues against.

Where your data lives

Launch, Team and Division run on shared infrastructure. Your data is stored in a shared database, separated from other customers by enforced database-level access controls, and is never accessible to another customer. Your monitoring, credentials and incident history remain private to your organisation.

Enterprise runs in an AWS account created for your organisation, with its own database. No infrastructure is shared with any other customer, and you choose the region.

Growing past a tier

Your tier is the number of services it covers.

The limit is the tier you chose

Each tier covers a number of services and the instance holds you to it: at your limit, onboarding another asks you to upgrade first. Nothing is throttled and nothing stops watching — services already onboarded keep running exactly as before.

You see the number before you hit it

Your console shows how much of the allowance is used, on the page where you onboard. Upgrading takes effect immediately and the same instance carries on — no migration, no new hostname, nothing to move.

Questions

The ones that actually get asked.

What counts as a service?

A production service you want the fleet to watch and act on — typically something with its own deployment, its own golden signals and its own on-call expectation. Not every repository, and not every container. Only services you have actually finished onboarding are counted; a half-configured one is not billed.

How long does provisioning really take?

Minutes, not days. The deployment packs are already built and tested against live accounts, so provisioning runs something that works rather than assembling it on the day. If it fails, we roll back and tell you — and nothing is charged for an instance that never came up.

Can we run it in our own AWS account?

Yes, on Enterprise. Your instance runs in an account created for your organisation with its own database, and nothing is shared with another customer. For running entirely inside infrastructure you already own, that is our managed service rather than this product — worth a conversation about which fits.

What happens to our data if we leave?

You can export it at any time, and we delete it on request. Suspension for a late invoice does not delete anything — your data is retained and comes back when the payment clears.

Do you need production credentials?

You choose what to connect. Credentials are encrypted at rest with a key held by your instance, and the fleet operates at the autonomy level you set — starting at observe-and-recommend, where it changes nothing.

Is there a free trial?

Not a free one, but a cheap one, and there is one for each product. Thalamus SRE Staging is $250 for a week and ThalamusTrace Staging is $50 for a week — a real instance either way, with your own database and up to two services onboarded. SRE gives you the full agent fleet against your existing monitoring; ThalamusTrace gives you an OTLP endpoint to point an exporter at. Take one or both, depending on which question you are trying to answer.

At the end of the week it pauses rather than disappearing. Promote to any tier and the same instance carries on with your configuration and data intact, or stop and we tear it down. Nothing is deleted while you decide.

Every tier runs real infrastructure from the moment it exists, which is why there is no free version — a running instance costs us money whether or not you use it. Running both Staging instances together is the cheapest way to see the handoff between them: ThalamusTrace raises a correlated problem, and the SRE fleet picks it up and investigates.

Next step

Bring your incident numbers, not a shortlist.

A 45-minute look at how your estate actually runs — incident volume, MTTR, what recurs, where the hours go. If there is not enough repetitive volume to justify going further, we will say so.