AI Observability
Purpose
The AI Observability module is designed for observability of AI infrastructure in Smart Monitor.
It collects AI environment telemetry, normalizes it into a unified set of gen_ai_* indexes, and provides ready boxed content for operations.
The module covers the following observability directions:
- LLM gateway — requests, model routing, errors, latency
- Inference runtime — performance and load of inference service
- GPU environment — graphics card status: temperature, memory, load, errors
- AI agents — questions, answers, agent steps, tool calls, traces
- Local AI clients — monitoring Claude Code and Codex as telemetry sources
- LLM costs — tokens, spend data by models and providers
- Service availability — general operational slice of AI components
Module Components
- Dashboards — 8 dashboards for all AI observability directions
- Inventory — automatically maintained GPU asset
- Service Monitor Toolkit global metrics — 6 key operational metrics for end-to-end monitoring
- Data sources — unified set of
gen_ai_*indexes with normalized fields
Documentation Sections
Data Sources
Indexes, collection architecture, and normalized fields of the AI Observability module.
AI Observability Module Installation
3 items
AI Observability: Capability
3 items