QA metrics

Custom metrics, Looker reports, and bug taxonomies.

Custom QA Metrics

Custom metrics tracked across 30+ sprints at Questrade (2022–2024).

Tracked in Google Sheets + Jira; visualized via Grafana (ops dashboards) and Looker Studio (Quality Reports for stakeholders)

How metrics were visualized → QA Observability Platform

Quality Reports · Looker Studio

After the Grafana/Influx pipeline, I published sprint Quality Reports in Google Looker Studio for stakeholders: test inventory, automation share, pyramid coverage, bug ratio per ticket, PPI and sprint achievement.

373 Test cases tracked
27.9% Automated
5 Projects in pyramid view
5+ Consecutive sprint reports

Same lineage as the observability platform: spreadsheet → Grafana/InfluxDB → Looker executive reports. A related Customer Lifecycle Looker dashboard sat alongside the sprint report.

Task Score

In Q4 2024 at Questrade I applied Task Score — a delivery quality metric that evaluates, per task, whether the team followed quality gates before reaching production. The premise: pass/fail rates alone don't show how consistently processes are followed sprint to sprint. Task Score makes the invisible visible — and turns retrospectives from feelings into data.

Click the criteria to simulate a task score 30/30 Excellent
Questrade · Q4 2024 — real data
24Q4C1
28.67/30
Excellent
24Q4C2
23.43/30
Good
24Q4C3
27.80/30
Excellent
Q4 Average
26.6/30
Excellent
TaskScore = SIT_Deploy(5) + Pipeline(5) + QA_Review(5) + No_Major_Bug(5) + No_Simple_Rejection(5) + Bug_Count(5|3|0)
→ avg per cadence = AVERAGEIF(tasks, cadence_id, score_col)

Before Task Score, quality issues were caught reactively — after bugs reached staging or production. The metric surfaces process gaps per cadence, not just counts, enabling targeted retrospectives. The C2 dip to 23.43 (Good) was directly linked to two high-severity bugs in a complex integration — it prompted a mid-cadence adjustment that brought C3 back to 27.80. Task Score is now a leading quality indicator in sprint retrospectives at Questrade.

PPI · SRM · RBD
PPI

Ping Pong Index

Ratio of total bugs to total fix attempts. A value of 1.0 means every bug was fixed correctly on the first try — values below 1.0 reflect ping-pong cycles between dev and QA.

PPI = ΣBug / ΣPingPong Each bug's PingPong starts at 1 · +1 on every failed fix returned to dev · 1.0 = perfect (one-time fix for all bugs)

Each bug starts with PingPong = 1 · +1 per failed fix returned to dev. So ΣPingPong ≥ ΣBug, and PPI ≤ 1.00. With 0 bugs, PingPong is 0 (disabled) and PPI is 1.00 Excellent.

PPI = 10 / 10 = 1.00 Excellent
ΣBug
10
ΣPingPong
10
Signals bug report clarity — low PPI often means ambiguous descriptions, not just dev errors
A bug with PingPong = 1 is a clean, first-time fix — the ideal baseline for every ticket
Trending toward 1.0 over sprints is concrete evidence that dev–QA alignment is improving
Simulated PPI · Excellent Demo values — not a client sprint. Mean across 30+ sprints stayed in Excellent.
RBD

Real Bug Detection

Percentage of reported issues that are genuine defects. Captures the cost of false positives — time spent opening, analyzing, and closing tickets that weren't real bugs.

RBD = (ΣBug − NoRealBug) / ΣBug × 100% NoRealBug = issues closed as "Invalid" · 100% = every reported ticket was a real defect

NoRealBug ≤ ΣBug. With 0 bugs, NoRealBug stays at 0 (disabled) and RBD is 100% Excellent.

RBD = (10 − 0) / 10 = 100% Excellent
ΣBug
10
NoRealBug
0
Reflects QA's grasp of acceptance criteria — low RBD suggests requirements need more precision
Quantifies wasted dev effort: each false positive consumes time to investigate and close
Indicates whether QA has adequate time to test — rushed cycles produce more false positives
Simulated RBD · Excellent Demo values — not a client sprint. Higher % means fewer false positives.
SRM

Simple Rejection Metric

Count of tasks returned to dev before testing even begins — because too many acceptance criteria were unimplemented. A raw count per sprint; lower is better, 0 is the goal.

SRM = tasks rejected before testing begins (count per sprint) Rejection requires a formal comment listing missing ACs · tracked per developer and per sprint

Adjust rejected-task count for the sprint. 0 is Ideal · lower is better.

SRM = 0 rejected tasks Ideal
Rejected tasks / sprint
0
Forces AC alignment before work reaches QA — eliminating wasted test cycles at the source
Per-developer tracking identifies who needs clearer requirements or stricter pre-QA checklists
Consistently hitting 0 for multiple sprints is the strongest signal of a mature dev–QA handoff
Simulated SRM · Ideal Demo count — not a client sprint. Goal is sustained 0.
Metrics in study

Metrics I'm currently studying to apply in my day-to-day QA work at companies — extending the PPI · SRM · RBD framework with prevention, automation health, and shift-left signals. Charts below use sanitized illustrative data, not production metrics from a client.

Pipeline & release health

Candidate indicators I'm evaluating for sprint retrospectives and QA observability — same collection patterns I'd use in practice (Sheets + Jira + Grafana).

Test pyramid

The test pyramid shows how automated tests should be spread by granularity: many fast, isolated checks at the base (unit), broader contract and service checks in the middle (integration, component), and a small set of full user journeys at the top (E2E). A healthy pyramid catches most defects early — where feedback is cheapest — and keeps slow, brittle UI tests to a minimum.

Test pyramid Four tiers: Unit, Integration, Component, E2E. Counts shown in external labels. Unit Integration Component E2E 2,161 974 276 87
Unit
· xUnit
2,161
Integration
· Postman/Newman
974
Component
· Cypress CT
276
E2E
· Cypress + Playwright
87
3,498 tests
Candidate metrics catalog · draft Draft

Catalog of candidate metrics I want to define next (formula, bands, Jira/CI source). Not instrumented yet; each card is a design brief so the PPI · SRM · RBD pattern can expand without inventing vanity KPIs.

Handoff & prevention

Signals that explain SRM and Task Score: readiness before QA starts, not only defects after.

ACR

AC Clarity Rate

Share of stories whose acceptance criteria are testable before kickoff. Low ACR usually precedes Simple Rejects.

In study
ACR = ΣStoriesWithTestableAC / ΣStoriesCommitted × 100% Testable AC = given/when/then or equivalent checkable outcomes · Source: Jira + refinement checklist

Related: Explains SRM; pair with Task Score refinement gates

DOF

Definition of Done Failures

Count of tasks that failed Definition of Done at QA intake in the sprint. Binary complement to Task Score.

In study
DOF = ΣTasksFailedDoD (per sprint) Failed DoD = missing evidence, tests, or AC coverage at handoff · Source: Jira workflow + checklist

Related: Sibling of SRM and Task Score criteria 1-5

BCR

Blocked Cycle Ratio

Share of cycle time spent blocked waiting on AC, environment, or test data. Turns delay into a root-cause signal.

In study
BCR = ΣBlockedDays / ΣCycleDays × 100% Blocked = waiting AC, env, data, or dependency · Source: Jira time-in-status

Related: Pairs with TTQR; explains slow handoffs without blaming QA capacity

Defect cycle

Complements PPI and RBD: reopen quality, resolve speed by severity, and ticket noise.

TTR

Time to Resolve

Median time from bug opened to closed, sliced by severity. Links the severity taxonomy to delivery speed.

In study
TTR = median(ClosedAt − OpenedAt) by severity Track Blocker/Critical separately from Minor/Low · Source: Jira · Target bands are for Critical+

Related: Uses severity ladder; pairs with PPI when fixes bounce

ROR

Reopen Rate

Bugs reopened after close divided by bugs closed. PPI counts ping-pong while open; ROR catches premature closes.

In study
ROR = ΣReopened / ΣClosed × 100% Reopened = status returns to open/reopened after Done · Source: Jira transitions

Related: Natural mirror of PPI; same Jira bug set

DIR

Duplicate / Invalid Ratio

Noise tickets (invalid + duplicate) over all reported bugs. RBD focuses on real defects; DIR isolates ticket waste.

In study
DIR = (ΣInvalid + ΣDuplicate) / ΣBug × 100% Invalid + Duplicate resolutions only · Source: Jira resolution field

Related: Splits RBD noise into actionable ticket hygiene

Automation & CI

Raises Flakiness and TPI from sparklines to formulable signals: ROI, signal quality, recovery, and coverage lag.

ARR

Automation ROI Ratio

Hours saved by automation versus hours spent maintaining the suite. Makes the qametrics story measurable per sprint.

In study
ARR = HoursSaved / HoursMaintained Saved = manual cycles avoided · Maintained = authoring + flake fixes · Source: Sheets + CI timing

Related: Narrative link to qametrics (1-3h/sprint to minutes)

SFR

Signal-to-Flake Ratio

Real failures divided by real failures plus flakes. More actionable than flake% alone when deciding to trust CI.

In study
SFR = ΣRealFails / (ΣRealFails + ΣFlakes) Flake = failed then passed on retry without code change · Source: Cypress Cloud / CI

Related: Upgrades the Flakiness study card into a decision metric

MTTG

Mean Time to Green

Average wall time from a red CI run to the next green on the same pipeline. Cost of an unstable suite.

In study
MTTG = mean(GreenAt − RedAt) Per primary QA pipeline (E2E/API) · Source: GitHub Actions / Cypress Cloud

Related: Pairs with SFR; shows operational cost of flake

CLR

Coverage Lag Rate

Share of stories that shipped code without a new automated test in the same sprint. Shift-left without vanity coverage%.

In study
CLR = ΣStoriesWithoutNewTest / ΣStoriesWithCode × 100% Exclude pure docs/config · Source: Jira + PR test file diff

Related: Complements TPI; pressures same-sprint automation

Release & risk

Go/no-go and post-release cost: readiness composite, hotfix load, and change without test evidence.

RRI

Release Readiness Index

Composite 0-100 score from escapes, open Blocker/Critical, flake, and SRM. A numeric go/no-go aligned to senior QA ownership.

In study
RRI = weighted(DER, OpenSev, SFR, SRM) → 0-100 Weights tuned per product risk · Source: Jira + CI + DER sheet

Related: Rolls up DER, Flakiness, SRM, severity into one gate

HFR

Hotfix Frequency

Hotfixes divided by planned releases in the period. Complements DER with operational cost after escape.

In study
HFR = ΣHotfixes / ΣPlannedReleases Hotfix = unplanned prod fix train · Source: release calendar + Jira

Related: Pairs with DER; shows blast radius after escape

CRV

Change Risk Violations

Deploys (or CRs) that changed in-scope code without linked test evidence. Closes the loop on post-deploy verification.

In study
CRV = ΣDeploysWithoutEvidence / ΣDeploys × 100% Evidence = automated run or signed manual pack on changed scope · Source: CR + CI artifacts

Related: Formalizes CR verification / go-no-go discipline

Regulated domain

KYC/CLP-specific signals that generic escape rate cannot tell: compliance escapes and IDV/MFA gate reliability.

CDE

Compliance Defect Escape

PROD escapes limited to regulated flows (KYC, tax, MFA, IDV). DER for the compliance slice only.

In study
CDE = ΣComplianceEscapes / ΣComplianceChanges × 100% Tag bugs with compliance=true · Source: Jira labels + release notes

Related: Domain-specific DER for KYC/CLP storytelling

IGR

IDV / Gate Reliability

Correct approve/reject outcomes for IDV and MFA gates in SIT versus expected fixtures (Mitek, TOTP, document paths).

In study
IGR = ΣCorrectGateOutcomes / ΣGateAttempts × 100% Includes selfie/doc approve+reject and MFA success/fail paths · Source: SIT runs + device lab

Related: Unique mobile+compliance metric from Mitek/MFA work

Coverage KPIs

Coverage with a target and a decision: change risk on the PR, critical paths in the domain, and AC-to-evidence linkage. Not raw line coverage %.

BRC

Branch Risk Coverage

Share of changed branches in a PR that gained a new or updated automated test. Shift-left coverage of the diff, not the whole repo.

KPI
BRC = ΣChangedBranchesWithTest / ΣChangedBranches × 100% KPI target ≥ 90% · Fail action: block merge or require waiver · Source: PR diff + CI coverage on changed files

Related: Complements CLR and CRV; merge gate for shift-left

CRC

Critical Path Coverage

Share of business-critical paths (KYC, tax, MFA, IDV) with current automated or signed evidence. Coverage that means go/no-go, not Sonar line %.

KPI
CRC = ΣCriticalPathsCovered / ΣCriticalPaths × 100% KPI target = 100% on release train · Fail action: hold go/no-go · Source: path inventory + Zephyr/CI evidence

Related: Pairs with CDE and RRI for regulated release readiness

SCC

Spec-to-Case Coverage

Share of sprint acceptance criteria with linked automated run or signed manual pack. Turns AC clarity into measurable DoD evidence.

KPI
SCC = ΣACWithEvidence / ΣACCommitted × 100% KPI target ≥ 95% per sprint · Fail action: fail DoD / open ACR gap · Source: Jira AC + Zephyr / CI links

Related: Closes the loop with ACR, DOF, and Looker test-case tracking

Defect severity taxonomy

Bug severity describes the impact of a defect — on users, data integrity, and whether testing or release can continue — independent of priority (when the team schedules the fix). I classify every bug I file on a five-level scale (Blocker → Low), using the same taxonomy most teams adopt in Jira and similar tools. How I apply it in practice: I assign severity at filing time with reproducible evidence (environment, steps, affected layer); in triage I push back on inflated Blocker/Critical labels and on cosmetic issues logged too high; open Blocker, Critical, or Major defects on in-scope flows block release until resolved or explicitly accepted; in retrospectives I group reopen and escape trends by severity to expose process gaps rather than blame individuals. Severity is a shared language for risk — it keeps conversations about scope, readiness, and quality metrics grounded in impact, not urgency alone.

Bug severity levels Stepped ladder from Blocker through Critical, Major, Minor to Low Blocker Critical Major Minor Low
Severity ladder · highest impact left → lowest right
Blocker

Stops all testing or release. No workaround. Production or SIT unusable for the affected flow — e.g. KYC onboarding blocked, auth down, data corruption preventing account updates.

Critical

Core feature failure with no acceptable workaround. System crash, data loss, or regulated path broken — MFA/IDV rejection loop, incorrect tax residency persisted, API returning 500 on mandatory steps.

Major

Significant defect in a feature; system partially usable or workaround exists. Wrong validation on a secondary field, broken edge case in Annual Review, flaky E2E on non-blocking step with manual fallback.

Minor

Small issue in a non-critical area. Does not block completion of the user journey — misaligned label, incorrect helper text, minor API field mismatch with no compliance impact.

Low

Cosmetic or copy issue. Typos, spacing, colour contrast suggestions, non-blocking UI polish — logged for backlog refinement, not release gates.

Severity ≠ Priority. Severity is impact on users and release flow; priority is when the team schedules the fix. A Low-severity cosmetic bug can still be High priority before a launch — see Bug priority for how I separate the two.

Defect priority taxonomy

Bug priority answers when a defect should be fixed — relative to other work in the backlog and the current sprint — independent of severity (how bad the impact is). I use a five-level scale (Highest → Lowest), aligned with common Jira priority fields. How I apply it in practice: at filing I suggest an initial priority based on sprint goals but leave final ranking to triage with PM and dev lead; I escalate to Highest when a fix must land in the current sprint or blocks another team; Medium is the default queue for defects that should ship soon but can wait for capacity; Low and Lowest go to backlog refinement unless a release date or dependency forces them up; in planning I map open Highest/High items to sprint capacity so severity-heavy bugs are not starved by urgency alone. Priority is a scheduling tool — it keeps the board honest about what ships next without confusing urgency with impact.

Bug priority levels Stepped ladder from Highest through High, Medium, Low to Lowest Highest High Medium Low Lowest
Priority ladder · fix first left → defer right
Highest

Must be resolved in the current sprint or immediately — blocks release, another squad, or a hard dependency. Often paired with high severity but driven by schedule, not impact alone.

High

Top of the backlog — target next sprint or hotfix window. Important for an upcoming milestone, regulatory date, or demo; team should pull before Medium items when capacity allows.

Medium

Default queue — should be fixed in a reasonable timeframe but can wait for sprint planning. Most functional defects land here after triage unless urgency or impact pushes them up or down.

Low

Fix when capacity allows — no sprint commitment. Acceptable to defer across releases if higher-priority work fills the board; revisit in backlog refinement.

Lowest

Nice-to-have — park until explicitly prioritized. Often cosmetic or edge-case items; may never ship unless bundled with related work or a polish sprint.

Pair severity + priority. Example: a typo on a legal disclaimer can be Low severity but High priority before go-live; a Critical-severity defect in a dormant feature might stay Medium priority until that feature is re-enabled. I document both fields so triage stays transparent. Nine prioritization frameworks →