Platform

See Everything. Understand Everything.

ForgeCrux One Observability is the telemetry plane for APIs, LLMs, MCP tools, and agents. Platform, SRE, security, and FinOps teams ingest logs, metrics, and traces from every gateway hop, then act from one dashboard—maps, alerts, incident tools, and audit replay included.

4

estates in one plane

AI, MCP, agent, and API telemetry with a shared trace ID.

OTEL

collector standard

OpenTelemetry, ForgeCrux agents, and auto-instrumented SDKs.

SLO

alerts that matter

Latency, errors, tokens, cost spikes, and policy violations.

replay

audit trail

Session and config replay for SRE, security, and GRC.

Reference architecture

ForgeCrux One Observability

One telemetry plane for the entire AI, MCP, agent, and API estate — from collector install to dashboards, maps, incidents, and audit replay.

ForgeCrux One Observability Platform

Entire AI, MCP, agent & API estate

ForgeCruxProbing Deeper, Stacking Precision

Data installation & setup guide

  1. 1

    Deploy collectors

    Install ForgeCrux agents, auto-instrument SDKs, or configure OpenTelemetry collectors on infrastructure and in applications.

  2. 2

    Connect estate sources

    Authenticate and link AI, MCP, agent, and API data sources via the console or API.

  3. 3

    Configure data pipelines

    Define data flows, parsing, filtering, and reduction rules.

  4. 4

    Validate & auto-discover

    Confirm data is arriving, and view auto-generated maps of the estate and services.

Entire AI estate

Models, training, inference, vector stores

MCP estate

Model control plane — cross-cloud model management

Agent estate

Autonomous agents, task execution, multi-agent coordination

API estate

Gateways, REST/GraphQL, third-party integrations

ForgeCrux One Observability Core

Telemetry ingestion

Logs, metrics, traces

Contextual linking & distributed tracing

App → gateway → model → tool

AI/ML powered anomaly detection

Cost, latency, quality, policy

Real-time dashboards & alerts

SLOs, tokens, errors

User session & audit trail replay

Who, what, why

Visualization & action layers

Unified single pane of glass

Dashboard / CLI

Service maps & dependency view

Auto-discovered estate

Incident management integrations

PagerDuty, ServiceNow

Performance optimization insights

Cost, cache, routing

One trace across four gateways

Follow a user or agent task through API, AI, MCP, and Agent hops without stitching four APM products.

Cost, quality, and risk in one view

Token spend, model latency, tool errors, and policy blocks share the same timeline as API SLOs.

Install collectors, then discover

Deploy agents or OTEL, connect estate sources, configure pipelines, and auto-generate service maps.

Key Capabilities

Unified traces: app → API → agent → model → MCP → system of record
Logs, metrics, and traces from all four gateways
Token, cost, latency, error, and cache analytics
OpenTelemetry collectors, SDKs, and ForgeCrux agents
Service maps and auto-discovery of the estate
Anomaly detection on spend, quality, and policy hits
Real-time dashboards, SLOs, and burn-rate alerts
Session and audit trail replay for GRC
PagerDuty, ServiceNow, Slack, and SIEM integrations
Export to Grafana, Datadog, Prometheus, and Splunk

Complete Observability capabilities

Everything required to publish, secure, mediate, observe, and operate observability workloads on ForgeCrux.

Enterprise uses

Who runs ForgeCrux One Observability and what they measure.

  • SRE: SLOs, error budgets, and dependency maps across API and AI hops
  • Platform engineering: one telemetry contract for every gateway data plane
  • AI CoE: token spend, quality drift, and model latency by team and app
  • Agent operations: step-level traces, tool failures, and HITL wait time
  • Security: policy hits, prompt leakage signals, and immutable audit replay
  • FinOps: chargeback on tokens, MCP calls, and API products
  • Support: session replay from user request to backend without log spelunking
  • GRC: evidence packs for SOC 2, ISO 27001, GDPR, and HIPAA reviews

Installation & deployment

Collect from anywhere the estate already runs.

  • ForgeCrux observability agents next to API, AI, MCP, and Agent gateways
  • OpenTelemetry Collector as DaemonSet, sidecar, or gateway
  • Auto-instrumented SDKs for Python, TypeScript, Java, and Go
  • SaaS telemetry backend, hybrid (private ingest + SaaS UI), or self-hosted
  • Helm, Terraform, and Kubernetes operator for collector and backend
  • Regional pin and data-residency for traces, logs, and session payloads
  • Sampling, tail-based sampling, and reduction rules at the edge
  • High-availability ingest with disk buffer and retry

Setup & onboarding

Four steps from empty cluster to auto-discovered maps.

  • Step 1 — Deploy collectors: agents, SDKs, or OTEL on infra and apps
  • Step 2 — Connect estate sources: authenticate AI, MCP, agent, and API streams
  • Step 3 — Configure pipelines: parse, filter, redact, and reduce
  • Step 4 — Validate & auto-discover: confirm ingest and generate service maps
  • Attach resource attributes: org, environment, product, and cost center
  • Turn on default SLO packs for latency, errors, tokens, and policy blocks
  • Route alerts to PagerDuty, ServiceNow, Slack, or email
  • Export copies to Grafana, Datadog, Prometheus, or SIEM if required

Security & privacy

Telemetry is production data—treat it with the same controls as the gateways.

  • mTLS and workload identity from collectors to ingest
  • PII/PHI redaction and payload truncation in pipelines
  • RBAC on dashboards, traces, session replay, and export jobs
  • SSO for operators; break-glass for incident commanders
  • Immutable audit of who viewed a trace or replayed a session
  • Retention, legal hold, and residency policies per environment
  • Secrets never logged; credential scans on log bodies
  • Private link / VPC ingest for hybrid and self-hosted estates

Pipelines, maps & action

From raw signals to incidents and optimization.

  • Contextual linking: one trace ID from client through every gateway
  • AI/ML anomaly detection on cost, quality, latency, and error shape
  • Real-time dashboards for traffic, tokens, agents, and MCP tools
  • Service maps and dependency views auto-built from traces
  • Incident integrations: PagerDuty, ServiceNow, Jira, Slack
  • Performance insights: cache hit ratio, model routing, and slow policies
  • User session and audit trail replay for support and GRC
  • CLI and API for the same queries the dashboard uses

Enterprise architecture

Observability reference architecture

Enterprise data path for Observability: producers, ForgeCrux control, gateway enforcement, and systems of record.

Sources

AI estate

Models, vectors, training

MCP estate

Tools and servers

Agent estate

Runs, A2A, HITL

API estate

Proxies and products

Collect

Agents

ForgeCrux collectors

OTEL

SDKs and collector

Pipelines

Parse, filter, redact

Ingest

Logs, metrics, traces

Core

Linking

Distributed trace

Anomaly

AI/ML detectors

Dashboards

SLOs and cost

Replay

Session and audit

Action

Maps

Dependencies

Incidents

PagerDuty, SNOW

Optimize

Routing and cache

Export

Grafana, SIEM

Data flows

How requests, policies, and telemetry move through ForgeCrux in this solution.

Collector to dashboard

How a hop becomes a trace you can act on.

1

Instrument

Agent / SDK / OTEL

2

Pipeline

Parse • redact • sample

3

Observability core

Ingest • link • store

4

Detect

SLO • anomaly • policy

5

Act

Alert • ticket • replay

Cross-estate trace

One request spanning all four gateways.

1

Client

App or agent

2

API Gateway

Auth • quota

3

Agent + AI

Plan • model

4

MCP

Tool call

5

Map + replay

Full hop timeline

How teams run Observability on ForgeCrux

Deploy collectors

Roll out ForgeCrux agents or OpenTelemetry on gateways and apps. Start with one environment.

Connect estate sources

Authenticate AI, MCP, agent, and API streams so they share a trace context and org labels.

Configure pipelines

Add parse, PII redaction, sampling, and cost-center attributes before production volume arrives.

Validate and auto-discover

Confirm ingest health, generate service maps, and attach default SLO packs.

Wire incidents

Send burn-rate and anomaly alerts to PagerDuty or ServiceNow with the trace attached.

Prove and optimize

Use replay for GRC, then routing and cache insights to cut latency and token spend.

Ready to get started with Observability?

Talk to our team about deploying Observability in your enterprise environment.