Platform
See Everything. Understand Everything.
4
estates in one plane
AI, MCP, agent, and API telemetry with a shared trace ID.
OTEL
collector standard
OpenTelemetry, ForgeCrux agents, and auto-instrumented SDKs.
SLO
alerts that matter
Latency, errors, tokens, cost spikes, and policy violations.
replay
audit trail
Session and config replay for SRE, security, and GRC.
Reference architecture
ForgeCrux One Observability
One telemetry plane for the entire AI, MCP, agent, and API estate — from collector install to dashboards, maps, incidents, and audit replay.
ForgeCrux One Observability Platform
Entire AI, MCP, agent & API estate
Data installation & setup guide
- 1
Deploy collectors
Install ForgeCrux agents, auto-instrument SDKs, or configure OpenTelemetry collectors on infrastructure and in applications.
- 2
Connect estate sources
Authenticate and link AI, MCP, agent, and API data sources via the console or API.
- 3
Configure data pipelines
Define data flows, parsing, filtering, and reduction rules.
- 4
Validate & auto-discover
Confirm data is arriving, and view auto-generated maps of the estate and services.
Entire AI estate
Models, training, inference, vector stores
MCP estate
Model control plane — cross-cloud model management
Agent estate
Autonomous agents, task execution, multi-agent coordination
API estate
Gateways, REST/GraphQL, third-party integrations
ForgeCrux One Observability Core
Telemetry ingestion
Logs, metrics, traces
Contextual linking & distributed tracing
App → gateway → model → tool
AI/ML powered anomaly detection
Cost, latency, quality, policy
Real-time dashboards & alerts
SLOs, tokens, errors
User session & audit trail replay
Who, what, why
Visualization & action layers
Unified single pane of glass
Dashboard / CLI
Service maps & dependency view
Auto-discovered estate
Incident management integrations
PagerDuty, ServiceNow
Performance optimization insights
Cost, cache, routing
One trace across four gateways
Follow a user or agent task through API, AI, MCP, and Agent hops without stitching four APM products.
Cost, quality, and risk in one view
Token spend, model latency, tool errors, and policy blocks share the same timeline as API SLOs.
Install collectors, then discover
Deploy agents or OTEL, connect estate sources, configure pipelines, and auto-generate service maps.
Key Capabilities
Complete Observability capabilities
Everything required to publish, secure, mediate, observe, and operate observability workloads on ForgeCrux.
Enterprise uses
Who runs ForgeCrux One Observability and what they measure.
- SRE: SLOs, error budgets, and dependency maps across API and AI hops
- Platform engineering: one telemetry contract for every gateway data plane
- AI CoE: token spend, quality drift, and model latency by team and app
- Agent operations: step-level traces, tool failures, and HITL wait time
- Security: policy hits, prompt leakage signals, and immutable audit replay
- FinOps: chargeback on tokens, MCP calls, and API products
- Support: session replay from user request to backend without log spelunking
- GRC: evidence packs for SOC 2, ISO 27001, GDPR, and HIPAA reviews
Installation & deployment
Collect from anywhere the estate already runs.
- ForgeCrux observability agents next to API, AI, MCP, and Agent gateways
- OpenTelemetry Collector as DaemonSet, sidecar, or gateway
- Auto-instrumented SDKs for Python, TypeScript, Java, and Go
- SaaS telemetry backend, hybrid (private ingest + SaaS UI), or self-hosted
- Helm, Terraform, and Kubernetes operator for collector and backend
- Regional pin and data-residency for traces, logs, and session payloads
- Sampling, tail-based sampling, and reduction rules at the edge
- High-availability ingest with disk buffer and retry
Setup & onboarding
Four steps from empty cluster to auto-discovered maps.
- Step 1 — Deploy collectors: agents, SDKs, or OTEL on infra and apps
- Step 2 — Connect estate sources: authenticate AI, MCP, agent, and API streams
- Step 3 — Configure pipelines: parse, filter, redact, and reduce
- Step 4 — Validate & auto-discover: confirm ingest and generate service maps
- Attach resource attributes: org, environment, product, and cost center
- Turn on default SLO packs for latency, errors, tokens, and policy blocks
- Route alerts to PagerDuty, ServiceNow, Slack, or email
- Export copies to Grafana, Datadog, Prometheus, or SIEM if required
Security & privacy
Telemetry is production data—treat it with the same controls as the gateways.
- mTLS and workload identity from collectors to ingest
- PII/PHI redaction and payload truncation in pipelines
- RBAC on dashboards, traces, session replay, and export jobs
- SSO for operators; break-glass for incident commanders
- Immutable audit of who viewed a trace or replayed a session
- Retention, legal hold, and residency policies per environment
- Secrets never logged; credential scans on log bodies
- Private link / VPC ingest for hybrid and self-hosted estates
Pipelines, maps & action
From raw signals to incidents and optimization.
- Contextual linking: one trace ID from client through every gateway
- AI/ML anomaly detection on cost, quality, latency, and error shape
- Real-time dashboards for traffic, tokens, agents, and MCP tools
- Service maps and dependency views auto-built from traces
- Incident integrations: PagerDuty, ServiceNow, Jira, Slack
- Performance insights: cache hit ratio, model routing, and slow policies
- User session and audit trail replay for support and GRC
- CLI and API for the same queries the dashboard uses
Enterprise architecture
Observability reference architecture
Enterprise data path for Observability: producers, ForgeCrux control, gateway enforcement, and systems of record.
Sources
AI estate
Models, vectors, training
MCP estate
Tools and servers
Agent estate
Runs, A2A, HITL
API estate
Proxies and products
Collect
Agents
ForgeCrux collectors
OTEL
SDKs and collector
Pipelines
Parse, filter, redact
Ingest
Logs, metrics, traces
Core
Linking
Distributed trace
Anomaly
AI/ML detectors
Dashboards
SLOs and cost
Replay
Session and audit
Action
Maps
Dependencies
Incidents
PagerDuty, SNOW
Optimize
Routing and cache
Export
Grafana, SIEM
Data flows
How requests, policies, and telemetry move through ForgeCrux in this solution.
Collector to dashboard
How a hop becomes a trace you can act on.
Instrument
Agent / SDK / OTEL
Pipeline
Parse • redact • sample
Observability core
Ingest • link • store
Detect
SLO • anomaly • policy
Act
Alert • ticket • replay
Cross-estate trace
One request spanning all four gateways.
Client
App or agent
API Gateway
Auth • quota
Agent + AI
Plan • model
MCP
Tool call
Map + replay
Full hop timeline
How teams run Observability on ForgeCrux
Deploy collectors
Roll out ForgeCrux agents or OpenTelemetry on gateways and apps. Start with one environment.
Connect estate sources
Authenticate AI, MCP, agent, and API streams so they share a trace context and org labels.
Configure pipelines
Add parse, PII redaction, sampling, and cost-center attributes before production volume arrives.
Validate and auto-discover
Confirm ingest health, generate service maps, and attach default SLO packs.
Wire incidents
Send burn-rate and anomaly alerts to PagerDuty or ServiceNow with the trace attached.
Prove and optimize
Use replay for GRC, then routing and cache insights to cut latency and token spend.
Related Products
Control Plane
ForgeCrux One is the enterprise control plane for APIs, AI models, MCP tools, and agents. Platform, security, and SRE teams use it to install, configure, govern, and operate every gateway from one catalog, one policy engine, and one audit trail—across SaaS, hybrid, and air-gapped footprints.
AI Gateway
ForgeCrux AI Gateway is the single endpoint for multi-model access, intelligent routing, prompt control, guardrails, token and cost management, evaluation, and LLM observability—across OpenAI, Anthropic, Gemini, Bedrock, Azure OpenAI, and self-hosted models.
Agent Gateway
ForgeCrux Agent Gateway gives every agent an identity, permissions, routing, memory controls, guardrails, tracing, evaluation, cost limits, and lifecycle—so multi-agent systems can reach models, APIs, MCP tools, and data without unmanaged autonomy.
API Gateway
ForgeCrux API Gateway covers the full API lifecycle: proxies, products, policies, developer portal, analytics, monetization, hybrid runtime, and promotion across environments—on the same control plane as AI, MCP, and agents.
Ready to get started with Observability?
Talk to our team about deploying Observability in your enterprise environment.