ThalamusTrace Playground
ThalamusTrace Playground

This is the real product,
running on real telemetry.

Not a mockup and not a video — the same observability and AIOps console our customers run, loaded with metrics, traces, logs and correlated problems from a live microservice estate. Explore anything you like; nothing you do here changes it.

17services in the platform
28SLOs with live error budgets
4signals — metrics, traces, logs, events
OTLPnative — no proprietary agent

Guided walkthroughs

Watch a problem form itself Concurrent anomalies across services get correlated into one Problem with a causal root cause and a blast radius — instead of five separate alerts arriving at 3am.
▶ Explore with walkthrough 4 min
Spend an error budget in real time An SLO's burn rate crossing its threshold, the budget draining, and the alert that fires before the objective is actually missed.
▶ Explore with walkthrough 3 min
Follow one request across nine services A single trace through the estate — where the latency actually went, which span failed, and the topology it travelled.
▶ Explore with walkthrough 3 min
See what an LLM call really cost Per-call tokens, latency and quality by model, drift detection, and A/B prompt comparison — the same treatment as any other dependency.
▶ Explore with walkthrough 4 min
Onboard a service, end to end From an OTLP endpoint to a service with SLOs, alerts and an owner — the path you'd follow on your own estate.
▶ Explore with walkthrough 5 min

Start here

What you'll see

Metrics, traces, logs, infrastructure and AI observability behind a single navigation — so “what broke, why, who's affected, and what to do next” is one investigation, not four tools.

ThalamusTrace home: fleet health scores, services needing attention, and capability cards.
The launchpad — fleet health at a glance, and the services that need attention first.

Problems, correlated and explained

One Problem per incident, with the root cause and blast radius on the row before you open anything.

  • Incident, capacity, AI and SLO problems in one queue
  • Filter by status, category, impact and severity
Problems view: model drift, disk pressure, breaching SLOs, a service down and a capacity risk, with a problems-over-time chart.

Objectives and error budgets

Burn-rate alerting before the objective is missed, per service and per model.

SLOs view with objectives, current attainment and error-budget burn.

A topology built from traces

Dependencies discovered from real requests, not declared in a diagram.

Smartmap: a service topology derived from live traces.

AI and LLM observability

Cost, tokens, latency and quality per model, with per-call prompt and response explainability.

AI and LLM observability: cost, tokens, latency and quality by model.

Capacity, before it bites

Forecasts that raise a Problem ahead of a saturation breach rather than after it.

Capacity forecast for disk utilization, projecting a saturation breach.

Ground rules

Read-only, enforced by the server

Editing, saving and remediation are refused so the environment stays coherent for the next visitor. Blocked controls say so rather than disappearing.

No account, no install

Sessions last 24 hours and need no signup or card. Leave the tab open and come back to it.

The telemetry is real OTLP

An instrumented estate emitting genuine metrics, traces and logs, with failures injected on purpose — so the problems are ones the platform detected.

Run it on your own telemetry

ThalamusTrace is OTLP-native — point an existing exporter at it with OTEL_EXPORTER_OTLP_ENDPOINT. No proprietary agent, no re-instrumentation.

Talk to us ↗