Labsco
axiomhq logo

building-dashboards

β˜… 10

by axiomhq Β· part of axiomhq/skills

Designs and builds Axiom dashboards via API. Covers chart types, APL and metrics/MPL query patterns, SmartFilters, layout, and configuration options. Use when creating dashboards, migrating from Splunk, or configuring chart options.

πŸ”₯πŸ”₯πŸ”₯βœ“ VerifiedFreeQuick setup
🧩 One of 6 skills in the axiomhq/skills package β€” works on its own, and pairs well with its siblings.

This is the playbook your agent receives when the skill activates β€” you don't need to read it to use the skill, but it's here to audit before installing.

Building Dashboards

Philosophy

  1. Decisions first. Every panel answers a question that leads to an action.
  2. Overview β†’ drilldown β†’ evidence. Start broad, narrow on click/filter, end with raw logs.
  3. Rates and percentiles over averages. Averages hide problems; p95/p99 expose them.
  4. Simple beats dense. One question per panel. No chart junk.
  5. Validate with data. Never guess fieldsβ€”discover schema first.
  6. Compute what's asked, or defer. If a panel can't be computed, replace it with a Note documenting the blocker. Never substitute a different quantity, even disclosed. See Compute or Defer.

Entry Points

Starting fromWorkflow
Vague descriptionIntake β†’ check dataset kind β†’ design blueprint (APL or MPL) β†’ queries per panel β†’ deploy
TemplatePick template β†’ customize dataset/service/env β†’ deploy
Splunk dashboardExtract SPL β†’ translate via spl-to-apl β†’ map to chart types β†’ deploy
Grafana dashboardProject canonical panel spec (expr, legendFormat, unit, title, description) β†’ translate PromQL β†’ map chart types β†’ deploy. See reference/grafana-migration.md.
ExplorationUse axiom-sre to discover schema/signals β†’ productize into panels

Intake: What to Ask First

  1. Audience & decision

    • Oncall triage? (fast refresh, error-focused)
    • Team health? (daily trends, SLO tracking)
    • Exec reporting? (weekly summaries, high-level)
  2. Scope

    • Service, environment, region, cluster, endpoint?
    • Single service or cross-service view?
  3. Dataset kind. Run scripts/metrics/datasets <deploy> and check kind.

    • otel:metrics:v1 β†’ metrics dataset, follow the Metrics path.
    • anything else β†’ events/logs dataset, follow the APL path.

    Never run getschema on a metrics dataset. It returns 0 rows without error.

    APL path: discover fields with ['dataset'] | where _time between (ago(1h) .. now()) | getschema. Continue to steps 4–5.

    Metrics path:

    • scripts/metrics/metrics-spec <deploy> <dataset> β€” required before any MPL query.
    • scripts/metrics/metrics-info <deploy> <dataset> metrics | tags | tags <tag> values for discovery.
    • If discovery is empty, retry with --start 7 days ago (sparse metrics).
    • find-metrics <value> searches tag values, not metric names β€” use it only with a known entity name.
    • Skip to the Metrics/MPL Blueprint.
  4. Golden signals (APL path)

    • Traffic: requests/sec, events/min
    • Errors: error rate, 5xx count
    • Latency: p50, p95, p99 duration
    • Saturation: CPU, memory, queue depth, connections
  5. Drilldown dimensions (APL path)

    • What do users filter/group by? (service, route, status, pod, customer_id)

Dashboard Blueprint

Pick the blueprint matching the dataset kind.

APL Blueprint (events/logs datasets)

1. At-a-Glance (Statistic panels)

Single numbers that answer "is it broken right now?"

  • Error rate (last 5m)
  • p95 latency (last 5m)
  • Request rate (last 5m)
  • Active alerts (if applicable)

Time-based patterns that answer "what changed?"

  • Traffic over time
  • Error rate over time
  • Latency percentiles over time
  • Stacked by status/service for comparison

3. Breakdowns (Table/Pie panels)

Top-N analysis that answers "where should I look?"

  • Top 10 failing routes
  • Top 10 error messages
  • Worst pods by error rate
  • Request distribution by status

4. Evidence (LogStream + SmartFilter)

Raw events that answer "what exactly happened?"

  • LogStream filtered to errors
  • SmartFilter for service/env/route
  • Key fields projected for readability

Metrics/MPL Blueprint (metrics datasets)

Use align to $__interval using … for bucketing β€” $__interval is supplied by the dashboard runtime. Hard-coded windows over- or under-resolve. Validate every pipeline with scripts/metrics/mpl-validate-chart; both it and chart-add --mpl reject inline time ranges ([1h..]).

Exception: for sparse metrics where $__interval rounds to empty buckets, a fixed wider window (e.g. 1h) is acceptable; document why on the chart.

1. At-a-Glance (Statistic panels)

Current values β€” "what's the state right now?"

  • Use group using avg (gauges) or group using last (counters).
  • Read the metric's unit via metrics-info … metrics <m> info and pass it to chart-add --unit. Ratio metrics (0–1) need | map * 100 in MPL before --unit "%".

Trends over time β€” "what changed?"

  • align to $__interval using avg|sum|last.
  • Group by low-cardinality tags only (≀10 series per chart).
  • Embed the unit in --name ("P95 Latency (ms)", "Memory (MiB)"); scale magnitudes in MPL (| map / 1048576 for bytes β†’ MiB).

3. Breakdowns (TimeSeries or Table panels)

Per-entity detail β€” "where should I look?"

  • Metrics broken down by entity (host, pod, service).
  • Filter to keep series count manageable.
  • One dimension per panel; don't overload a single chart.

4. Entity State (TimeSeries or Table panels)

Boolean/state metrics β€” answer "what is on/off/active?"

  • Use align to $__interval using last.
  • Sparse state metrics may need a fixed wider interval (1h+).

Required Chart Structure

Each chart needs a unique kebab-case id (error-rate, p95-latency); every layout i must match one. Pass the same id to chart-add --id and layout-pack <id>:…. dashboard-assemble cross-checks before emit.


Compute or Defer

Each panel either computes the requested quantity, or it's replaced by a Note documenting the blocker. Substituting a different quantity is never acceptable β€” disclaimers don't reach whoever acts on the number.

Defer template (use chart-add --type Note):

**Deferred β€” blocked by:** <one-line reason>.

**Original spec:** <what the panel should compute, dimensions, unit>.

**To unblock:** <pointer to the fix>.

Common blockers: MPL parser limits, missing tag with no reverse-tag equivalent, missing metric with no OTel rename match. Full rationale: reference/design-playbook.md Β§ Substituting a Different Quantity.


Chart Types

TypeWhenKey constraint
StatisticSingle KPI, current valueQuery must return one row.
TimeSeriesTrends over time, percentile overlaysbin_auto(_time); percentiles_array() for multi-percentile.
TableTop-N lists, breakdownsBound with top N; control columns via project.
PieShare-of-total for ≀6 categoriesAggregate to ≀6 slices; never high-cardinality.
LogStreamRaw event inspectiontake 100–500; project-keep to relevant fields; filter hard.
HeatmapDistribution / latency densitysummarize histogram(field, buckets) by bin_auto(_time).
Scatter PlotCorrelate two metrics per groupsummarize avg(x), avg(y) by group.
SmartFilterInteractive filter barEach panel query needs declare query_parameters. See reference/smartfilter.md.
Monitor ListMonitor status displayNo APL β€” select monitors in UI.
NoteMarkdown context, headers, runbook linkschart-add --type Note --text "<md>".

Per-type APL recipes: reference/chart-cookbook.md.


APL Patterns

Time Filtering

Dashboard chart queries inherit time from the picker β€” omit _time filters. Ad-hoc queries (Axiom Query tab, axiom-sre) need an explicit where _time between (ago(1h) .. now()).

Bin Size Selection

Use bin_auto(_time) β€” it adjusts to the dashboard time window. Manual bin(_time, …) is only justified for non-standard cases (e.g. matching an upstream batch interval); document why.

Cardinality Guardrails

Bound summarize … by … with top N or a filter. Unbounded grouping on high-cardinality fields (user_id, trace_id) blows up.

| summarize count() by route | top 10 by count_   // bounded
| summarize count() by user_id                    // unbounded β€” avoid

Field Escaping

Fields with dots need bracket notation:

| where ['kubernetes.pod.name'] == "frontend"

Fields with dots IN the name (not hierarchy) need escaping:

| where ['kubernetes.labels.app\\.kubernetes\\.io/name'] == "frontend"

Recipes

Traffic, error-rate, latency-percentile, and other golden-signal APL recipes: reference/chart-cookbook.md.


Layout Composition

layout-pack packs charts row-major into the 12-column grid using per-type defaults (Statistic 3Γ—3, TimeSeries 6Γ—4, Table 6Γ—5, LogStream 12Γ—6, Note 12Γ—2). Override with id:WxH when needed. Section blueprints: reference/layout-recipes.md. Naming and panel-ordering conventions: reference/design-playbook.md.


Dashboard Settings

Refresh Rate

dashboard-assemble --refresh oncall|team|exec (60/300/900s) or pass an explicit integer (β‰₯60). Short refresh + long time range = expensive queries; pick the longer end for exec/weekly boards.

Sharing

API tokens create shared dashboards only (owner: "X-AXIOM-EVERYONE"); private dashboards aren't supported. Per-user data visibility is still enforced by dataset permissions.

URL Time Range Parameters

?t_qr=24h (quick range), ?t_ts=...&t_te=... (custom), ?t_against=-1d (comparison)


Sibling Skill Integration

  • spl-to-apl β€” Splunk SPL β†’ APL (timechart β†’ TimeSeries, stats β†’ Statistic/Table). See reference/splunk-migration.md.
  • axiom-sre β€” schema discovery via getschema, baseline exploration.
  • query-metrics β€” metrics dataset/tag/value discovery; same scripts vendored under scripts/metrics/.

Templates

Compose with chart-add + layout-pack + dashboard-assemble. Pre-built templates remain under reference/templates/ (blank.json, service-overview.json, service-overview-with-filters.json, api-health.json) for legacy use; dashboard-from-template instantiates them but assumes specific field names (service, status, route, duration_ms) and needs sed-fixing. Prefer composition for new work.


Common Pitfalls

ProblemCauseSolution
getschema returns 0 rowsDataset is otel:metrics:v1Use scripts/metrics/metrics-info for metrics discovery.
Metrics discovery returns emptySparse metrics outside the 24h default windowRetry with --start 7 days ago.
404 from metrics API callsUsed scripts/axiom-api (dashboard) instead of scripts/metrics/axiom-apiUse scripts/metrics/axiom-api for /v1/query/*, /v1/datasets.
Statistic shows 1 instead of 100% for a 0–1 ratioPercent enum doesn't auto-multiply| map * 100 in MPL, then chart-add --unit "%".
OTel histogram chart shows nonsenseHistogram aligned as a scalarUse bucket … using interpolate_cumulative_histogram (or _delta per temporality). See promql-to-mpl.md Β§ Histogram translation.
Grafana migration filters/groups on the wrong subsetRead expr without description, or vice versaProject all five panel fields before authoring; see reference/grafana-migration.md.
PromQL metric name not foundSkipped OTel rename rulesDrop _total, decompose histograms, normalise units; validate with metrics-info. Labels need reverse-tag discovery. See grafana-migration.md Β§ Name Mapping.
MPL chart aggregates across a dimension PromQL filtered/grouped onDropped a selector or by(...) during translationEvery {label=…} β†’ where; every by(…) β†’ group by. See reference/promql-to-mpl.md.
Panel shipped a different quantity than askedSubstituted instead of deferringReplace with a Note documenting the blocker. See Compute or Defer.
403 "creating private dashboards"API tokens only create shared dashboardsLeave owner as dashboard-assemble's default (X-AXIOM-EVERYONE).

Reference

  • reference/chart-config.md β€” All chart configuration options (JSON)
  • reference/metrics-mpl.md β€” Metrics/MPL chart contract and discovery scripts
  • reference/smartfilter.md β€” SmartFilter/FilterBar full configuration
  • reference/chart-cookbook.md β€” APL patterns per chart type
  • reference/layout-recipes.md β€” Grid layouts and section blueprints
  • reference/splunk-migration.md β€” Splunk panel β†’ Axiom mapping
  • reference/grafana-migration.md β€” Grafana panel β†’ Axiom mapping (canonical-spec projection, PromQLβ†’MPL pointers, OTel rename rules)
  • reference/promql-to-mpl.md β€” PromQL β†’ MPL translation rules (selectors, groupings, rate, histograms, ratios, reverse-tag discovery)
  • reference/design-playbook.md β€” Decision-first design principles
  • reference/templates/ β€” Ready-to-use dashboard JSON files

For APL syntax: https://axiom.co/docs/apl/introduction