Push Grafana Cloud dashboards and alert rules (dashboard- and alerts-as-code). Runs playbooks/deploy_dashboard.yml and playbooks/deploy_alerts.yml. TRIGGER when: the user wants to push, deploy, or update the Grafana dashboard or alert rules, apply dashboard-as-code, set up the cluster health dashboard, add or change a Grafana alert, or asks why no alerts exist. SKIP: if the user only wants to query Grafana Cloud (that is the MCP server from /sales-demos-mcp) or deploy Alloy (that is /sales-de...
Scanned 10/6/2026
npx -y skills add ericcames/sales.demos --skill sales-demos-dashboard --agent claude-codeInstalls into .claude/skills of the current project.
Are you the author of Sales Demos Dashboard?
Add the live security badge to your README — it updates automatically with every re-scan.
[](https://www.skillsdirectory.com/skills/ericcames-sales-demos-dashboard)More formats (shields.io, HTML) on the badges page. Keep it an A: scan every change in CI with Pro.
---
name: sales-demos-dashboard
description: "Push Grafana Cloud dashboards and alert rules (dashboard- and alerts-as-code). Runs playbooks/deploy_dashboard.yml and playbooks/deploy_alerts.yml. TRIGGER when: the user wants to push, deploy, or update the Grafana dashboard or alert rules, apply dashboard-as-code, set up the cluster health dashboard, add or change a Grafana alert, or asks why no alerts exist. SKIP: if the user only wants to query Grafana Cloud (that is the MCP server from /sales-demos-mcp) or deploy Alloy (that is /sales-demos-alloy)."
---
# sales-demos-dashboard
Push Grafana Cloud dashboards and alert rules defined as committed JSON. Issues
[#275](https://github.com/ericcames/sales.demos/issues/275) (dashboard) and
[#629](https://github.com/ericcames/sales.demos/issues/629) (alerts).
## There is an AAP path now too (#318)
`AAP Observability - 2 Deploy Dashboards` runs the same playbook from AAP, so
this no longer has to come off a laptop. Use whichever suits; the skill is still
the quicker loop while iterating on dashboard JSON.
**It is not per-environment, and that surprises people.** One Grafana Cloud
serves both environments, so the template exists in both controllers and pushes
to the *same* folder — running it from demo also updates what sandbox sees.
This skill contains **no logic**. All the work is in
[`playbooks/deploy_dashboard.yml`](../../../playbooks/deploy_dashboard.yml). See
`CLAUDE.md` → *Skills and playbooks*.
## What it does
1. Creates a "Sales Demos" folder in Grafana Cloud (idempotent)
2. Reads `playbooks/files/grafana/cluster-health.json`
3. Pushes the dashboard via the Grafana HTTP API with `overwrite: true`
The dashboard covers cluster nodes, KubeVirt VMs, AAP platform health, and
logs. A `cluster` template variable makes it work for both sandbox and demo.
`deploy_alerts.yml` does the same for alert rules
(`playbooks/files/grafana/alert-rules.json`): it PUTs one rule group,
`sales-demos-health`, into the same folder. The PUT replaces the whole group, so
a rule removed from the JSON is removed from Grafana. Seven rules, each labelled
by `cluster`:
| Rule | Fires when |
|---|---|
| Alloy federation down | a cluster reported metrics in the last hour but not now (5m) |
| AAP controller metrics down | the AAP metrics scrape answered in the last hour but not now (5m) |
| Running VM count dropped | fewer VMs running than 10 minutes ago (1m) — expected after a teardown |
| AAP jobs stuck pending | any pending job for 15m |
| Free-tier series budget above 80% | over 8,000 active series stack-wide (15m) |
| Node under disk pressure | kubelet reports DiskPressure on any node (immediate) — it is already evicting, and AAP job pods request no ephemeral-storage so they rank among the first taken (#782) |
| Node disk approaching the eviction threshold | a node's `/var` is above 82% used for 15m — eviction begins at 85% (#782) |
**No contact point is configured.** Firing alerts follow the stack's default
notification policy; adding a receiver would put an address in a public repo.
Rules stay editable in the UI (`X-Disable-Provenance`), and the next run puts the
committed version back.
AAP path: `AAP Observability - 3 Deploy Alerts`, not per-environment, like template 2.
## Preflight Check
Run these before doing anything else. Every one must pass.
```bash
VAULT_ID="sales.demos@$HOME/secrets/.vault_pass_sales_demos"
# 1. Vault password file exists
test -s "$HOME/secrets/.vault_pass_sales_demos" \
&& echo "pass: vault password file" \
|| echo "FAIL: ~/secrets/.vault_pass_sales_demos missing"
# 2. secrets.yml exists and is vault-encrypted
head -c 15 playbooks/group_vars/all/secrets.yml 2>/dev/null | grep -q '^\$ANSIBLE_VAULT' \
&& echo "pass: secrets.yml is vault-encrypted" \
|| echo "FAIL: secrets.yml missing or NOT encrypted — see /sales-demos-first-time"
# 3. Grafana Cloud Editor SA token is filled in
ansible-vault view playbooks/group_vars/all/secrets.yml --vault-id "$VAULT_ID" 2>/dev/null \
| python3 -c "
import sys, yaml
d = yaml.safe_load(sys.stdin) or {}
keys = ['grafana_cloud_url', 'grafana_cloud_editor_sa_token']
bad = [k for k in keys if d.get(k, 'CHANGEME') == 'CHANGEME' or k not in d]
print(('FAIL: missing or CHANGEME: ' + ', '.join(bad)) if bad
else 'pass: Grafana Cloud dashboard credentials filled in')
"
# 4. Dashboard JSON exists
test -f playbooks/files/grafana/cluster-health.json \
&& echo "pass: dashboard JSON exists" \
|| echo "FAIL: playbooks/files/grafana/cluster-health.json missing"
```
If any check fails, stop and tell the user exactly which one and the fix shown
beside it. Do not attempt the run with a failing prerequisite.
### If the Editor SA token is missing
The user must create it manually in Grafana Cloud:
1. Administration > Service Accounts > Add
2. Name: `sales-demos-editor`, Role: **Editor**
3. Add token > copy the `glsa_...` value
4. Add to vault as `grafana_cloud_editor_sa_token`
This is separate from the Viewer SA token used by the MCP server.
## Run
```bash
mkdir -p ~/ansible-logs
export ANSIBLE_LOG_PATH=~/ansible-logs/deploy-dashboard-$(date +%F-%H%M).log
./utilities/run-ansible.sh playbooks/deploy_dashboard.yml -i inventory \
--vault-id sales.demos@~/secrets/.vault_pass_sales_demos
```
**Always set `ANSIBLE_LOG_PATH`** — logs live outside the repo, in
`~/ansible-logs/`. Tell the user the path.
**No `--limit` needed.** This playbook targets localhost because Grafana Cloud
is a single external service. Do NOT pass `--limit sandbox` or `--limit demo`
— localhost is not in those groups and the play will skip with "no hosts
matched".
**`-i inventory` is required** even though the play targets localhost, because
Ansible needs the inventory path to resolve `group_vars/all/` for vault
variable loading.
This takes under 30 seconds.
### Alert rules
```bash
export ANSIBLE_LOG_PATH=~/ansible-logs/deploy-alerts-$(date +%F-%H%M).log
./utilities/run-ansible.sh playbooks/deploy_alerts.yml -i inventory \
--vault-id sales.demos@~/secrets/.vault_pass_sales_demos
# reversal — deletes the rule group, then asserts it is gone
./utilities/run-ansible.sh playbooks/deploy_alerts.yml -i inventory \
-e alerts_state=absent \
--vault-id sales.demos@~/secrets/.vault_pass_sales_demos
```
The playbook reads the group back and asserts every committed rule uid is there,
so a green run already means Grafana holds the rules. Confirm from the agent's
side anyway: `alerting_manage_rules` with `operation: list` should show the five
rules in folder "Sales Demos", each `normal` unless something is genuinely wrong.
## Verify via Grafana MCP
Use the Grafana MCP server to confirm the dashboard was pushed:
1. **Queries match:** `get_dashboard_panel_queries` with uid
`sales-demos-cluster-health` — returns every panel's query expression. This
is the fastest way to confirm a specific panel change landed.
2. **Dashboard exists:** `search_dashboards` with query `cluster health` —
should return "Sales Demos - Cluster Health" in the "Sales Demos" folder.
3. **Full model:** `get_dashboard_by_uid` with uid
`sales-demos-cluster-health` — returns the complete dashboard. Use
`get_dashboard_property` with a JSONPath to check a specific field without
pulling the whole model (e.g., `$.panels[*].options.textMode`).
Then open the dashboard URL printed by the playbook and confirm panels render
with live data.
## v1/v2 schema note
The playbook pushes via the legacy v1 API (`POST /api/dashboards/db`). This
Grafana Cloud stack stores dashboards in `v0alpha1` format
(`status.conversion.storedVersion`) and converts to v2 on read. The v1 write
path has been reliable for all fields so far, but if a future panel option
appears wrong in the live dashboard despite the v1 API returning the correct
value, check the v2 apiserver directly:
```
GET /apis/dashboard.grafana.app/v2/namespaces/stacks-<stack-id>/dashboards/sales-demos-cluster-health
```
`<stack-id>` is your stack's numeric ID: Grafana Cloud portal › your stack ›
Details, or the `stack_id` label on `grafanacloud_instance_info` in the
`grafanacloud-usage` data source. Like the stack URL, it identifies the
account, so it is never written into a tracked file (#635).
Compare `resourceVersion` and `generation` between the v2 response and what
the Grafana UI is rendering — a mismatch indicates read replica lag or a
stale client session, not a write failure.
## When it finishes
Report the dashboard URL and tell the user the dashboard is live in Grafana
Cloud. Remind them to select the `cluster` variable (e.g., `sandbox`) to see
data.
## If it fails
| Symptom | Cause | Fix |
|---|---|---|
| Assertion fails on `grafana_cloud_editor_sa_token` | Token not in vault | Create the Editor SA in Grafana Cloud UI, add to vault |
| `401 Unauthorized` | Token expired or revoked | Regenerate in Grafana Cloud > Service Accounts |
| `403 Forbidden` | Token has Viewer role, not Editor | Create a new SA with Editor role |
| `412 Precondition Failed` on folder creation | Folder already exists and was modified | Already handled by the playbook (accepts 200, 409, 412) |
| `Attempting to decrypt but no vault secrets found` | `--vault-id` missing | Add `--vault-id sales.demos@~/secrets/.vault_pass_sales_demos` |
| `no hosts matched` / skipping | `--limit` was passed | Remove `--limit` — this play targets localhost |
Never paste a Grafana Cloud URL or token into a commit message, issue, or PR.
This repo is public — see `CLAUDE.md`.
Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.
No comments yet. Be the first to comment!