AC Clarity Rate
Share of stories whose acceptance criteria are testable before kickoff. Low ACR usually precedes Simple Rejects.
Related: Explains SRM; pair with Task Score refinement gates
After the Grafana/Influx pipeline, I published sprint Quality Reports in Google Looker Studio for stakeholders: test inventory, automation share, pyramid coverage, bug ratio per ticket, PPI and sprint achievement.
Same lineage as the observability platform: spreadsheet → Grafana/InfluxDB → Looker executive reports. A related Customer Lifecycle Looker dashboard sat alongside the sprint report.
In Q4 2024 at Questrade I applied Task Score — a delivery quality metric that evaluates, per task, whether the team followed quality gates before reaching production. The premise: pass/fail rates alone don't show how consistently processes are followed sprint to sprint. Task Score makes the invisible visible — and turns retrospectives from feelings into data.
Before Task Score, quality issues were caught reactively — after bugs reached staging or production. The metric surfaces process gaps per cadence, not just counts, enabling targeted retrospectives. The C2 dip to 23.43 (Good) was directly linked to two high-severity bugs in a complex integration — it prompted a mid-cadence adjustment that brought C3 back to 27.80. Task Score is now a leading quality indicator in sprint retrospectives at Questrade.
Sanitized sample — illustrative trend, not client data
Ratio of total bugs to total fix attempts. A value of 1.0 means every bug was fixed correctly on the first try — values below 1.0 reflect ping-pong cycles between dev and QA.
Each bug starts with PingPong = 1 · +1 per failed fix returned to dev. So ΣPingPong ≥ ΣBug, and PPI ≤ 1.00. With 0 bugs, PingPong is 0 (disabled) and PPI is 1.00 Excellent.
Percentage of reported issues that are genuine defects. Captures the cost of false positives — time spent opening, analyzing, and closing tickets that weren't real bugs.
NoRealBug ≤ ΣBug. With 0 bugs, NoRealBug stays at 0 (disabled) and RBD is 100% Excellent.
Count of tasks returned to dev before testing even begins — because too many acceptance criteria were unimplemented. A raw count per sprint; lower is better, 0 is the goal.
Adjust rejected-task count for the sprint. 0 is Ideal · lower is better.
Metrics I'm currently studying to apply in my day-to-day QA work at companies — extending the PPI · SRM · RBD framework with prevention, automation health, and shift-left signals. Charts below use sanitized illustrative data, not production metrics from a client.
Candidate indicators I'm evaluating for sprint retrospectives and QA observability — same collection patterns I'd use in practice (Sheets + Jira + Grafana).
Study metrics · illustrative trends for portfolio exploration
The test pyramid shows how automated tests should be spread by granularity: many fast, isolated checks at the base (unit), broader contract and service checks in the middle (integration, component), and a small set of full user journeys at the top (E2E). A healthy pyramid catches most defects early — where feedback is cheapest — and keeps slow, brittle UI tests to a minimum.
Catalog of candidate metrics I want to define next (formula, bands, Jira/CI source). Not instrumented yet; each card is a design brief so the PPI · SRM · RBD pattern can expand without inventing vanity KPIs.
Signals that explain SRM and Task Score: readiness before QA starts, not only defects after.
Share of stories whose acceptance criteria are testable before kickoff. Low ACR usually precedes Simple Rejects.
Related: Explains SRM; pair with Task Score refinement gates
Count of tasks that failed Definition of Done at QA intake in the sprint. Binary complement to Task Score.
Related: Sibling of SRM and Task Score criteria 1-5
Share of cycle time spent blocked waiting on AC, environment, or test data. Turns delay into a root-cause signal.
Related: Pairs with TTQR; explains slow handoffs without blaming QA capacity
Complements PPI and RBD: reopen quality, resolve speed by severity, and ticket noise.
Median time from bug opened to closed, sliced by severity. Links the severity taxonomy to delivery speed.
Related: Uses severity ladder; pairs with PPI when fixes bounce
Bugs reopened after close divided by bugs closed. PPI counts ping-pong while open; ROR catches premature closes.
Related: Natural mirror of PPI; same Jira bug set
Noise tickets (invalid + duplicate) over all reported bugs. RBD focuses on real defects; DIR isolates ticket waste.
Related: Splits RBD noise into actionable ticket hygiene
Raises Flakiness and TPI from sparklines to formulable signals: ROI, signal quality, recovery, and coverage lag.
Hours saved by automation versus hours spent maintaining the suite. Makes the qametrics story measurable per sprint.
Related: Narrative link to qametrics (1-3h/sprint to minutes)
Real failures divided by real failures plus flakes. More actionable than flake% alone when deciding to trust CI.
Related: Upgrades the Flakiness study card into a decision metric
Average wall time from a red CI run to the next green on the same pipeline. Cost of an unstable suite.
Related: Pairs with SFR; shows operational cost of flake
Share of stories that shipped code without a new automated test in the same sprint. Shift-left without vanity coverage%.
Related: Complements TPI; pressures same-sprint automation
Go/no-go and post-release cost: readiness composite, hotfix load, and change without test evidence.
Composite 0-100 score from escapes, open Blocker/Critical, flake, and SRM. A numeric go/no-go aligned to senior QA ownership.
Related: Rolls up DER, Flakiness, SRM, severity into one gate
Hotfixes divided by planned releases in the period. Complements DER with operational cost after escape.
Related: Pairs with DER; shows blast radius after escape
Deploys (or CRs) that changed in-scope code without linked test evidence. Closes the loop on post-deploy verification.
Related: Formalizes CR verification / go-no-go discipline
KYC/CLP-specific signals that generic escape rate cannot tell: compliance escapes and IDV/MFA gate reliability.
PROD escapes limited to regulated flows (KYC, tax, MFA, IDV). DER for the compliance slice only.
Related: Domain-specific DER for KYC/CLP storytelling
Correct approve/reject outcomes for IDV and MFA gates in SIT versus expected fixtures (Mitek, TOTP, document paths).
Related: Unique mobile+compliance metric from Mitek/MFA work
Coverage with a target and a decision: change risk on the PR, critical paths in the domain, and AC-to-evidence linkage. Not raw line coverage %.
Share of changed branches in a PR that gained a new or updated automated test. Shift-left coverage of the diff, not the whole repo.
Related: Complements CLR and CRV; merge gate for shift-left
Share of business-critical paths (KYC, tax, MFA, IDV) with current automated or signed evidence. Coverage that means go/no-go, not Sonar line %.
Related: Pairs with CDE and RRI for regulated release readiness
Share of sprint acceptance criteria with linked automated run or signed manual pack. Turns AC clarity into measurable DoD evidence.
Related: Closes the loop with ACR, DOF, and Looker test-case tracking
Catalog only · bands are draft proposals until collected on real sprints
Bug severity describes the impact of a defect — on users, data integrity, and whether testing or release can continue — independent of priority (when the team schedules the fix). I classify every bug I file on a five-level scale (Blocker → Low), using the same taxonomy most teams adopt in Jira and similar tools. How I apply it in practice: I assign severity at filing time with reproducible evidence (environment, steps, affected layer); in triage I push back on inflated Blocker/Critical labels and on cosmetic issues logged too high; open Blocker, Critical, or Major defects on in-scope flows block release until resolved or explicitly accepted; in retrospectives I group reopen and escape trends by severity to expose process gaps rather than blame individuals. Severity is a shared language for risk — it keeps conversations about scope, readiness, and quality metrics grounded in impact, not urgency alone.
Stops all testing or release. No workaround. Production or SIT unusable for the affected flow — e.g. KYC onboarding blocked, auth down, data corruption preventing account updates.
Core feature failure with no acceptable workaround. System crash, data loss, or regulated path broken — MFA/IDV rejection loop, incorrect tax residency persisted, API returning 500 on mandatory steps.
Significant defect in a feature; system partially usable or workaround exists. Wrong validation on a secondary field, broken edge case in Annual Review, flaky E2E on non-blocking step with manual fallback.
Small issue in a non-critical area. Does not block completion of the user journey — misaligned label, incorrect helper text, minor API field mismatch with no compliance impact.
Cosmetic or copy issue. Typos, spacing, colour contrast suggestions, non-blocking UI polish — logged for backlog refinement, not release gates.
Severity ≠ Priority. Severity is impact on users and release flow; priority is when the team schedules the fix. A Low-severity cosmetic bug can still be High priority before a launch — see Bug priority for how I separate the two.
Bug priority answers when a defect should be fixed — relative to other work in the backlog and the current sprint — independent of severity (how bad the impact is). I use a five-level scale (Highest → Lowest), aligned with common Jira priority fields. How I apply it in practice: at filing I suggest an initial priority based on sprint goals but leave final ranking to triage with PM and dev lead; I escalate to Highest when a fix must land in the current sprint or blocks another team; Medium is the default queue for defects that should ship soon but can wait for capacity; Low and Lowest go to backlog refinement unless a release date or dependency forces them up; in planning I map open Highest/High items to sprint capacity so severity-heavy bugs are not starved by urgency alone. Priority is a scheduling tool — it keeps the board honest about what ships next without confusing urgency with impact.
Must be resolved in the current sprint or immediately — blocks release, another squad, or a hard dependency. Often paired with high severity but driven by schedule, not impact alone.
Top of the backlog — target next sprint or hotfix window. Important for an upcoming milestone, regulatory date, or demo; team should pull before Medium items when capacity allows.
Default queue — should be fixed in a reasonable timeframe but can wait for sprint planning. Most functional defects land here after triage unless urgency or impact pushes them up or down.
Fix when capacity allows — no sprint commitment. Acceptable to defer across releases if higher-priority work fills the board; revisit in backlog refinement.
Nice-to-have — park until explicitly prioritized. Often cosmetic or edge-case items; may never ship unless bundled with related work or a polish sprint.
Pair severity + priority. Example: a typo on a legal disclaimer can be Low severity but High priority before go-live; a Critical-severity defect in a dormant feature might stay Medium priority until that feature is re-enabled. I document both fields so triage stays transparent. Nine prioritization frameworks →