36 Research-Validated Signals That Predict Deployment Failures
The complete signal library for deploy risk scoring — organized by category with peer-reviewed citations from two decades of software defect prediction research.
Research foundation
These signals are drawn from two decades of Just-In-Time (JIT) software defect prediction research — academic studies from Microsoft Research, Google SRE, CMU, and peer-reviewed conferences including ICSE, FSE, MSR, and TSE. AUC (Area Under the ROC Curve) scores reflect predictive power on held-out deployment outcome datasets: 0.5 = random prediction; 1.0 = perfect prediction. Scores above 0.65 are considered practically significant in the defect prediction literature.
All 36 Signals
Research-validated signals spanning change complexity, author expertise, test coverage, schema changes, deployment context, historical patterns, and emerging AI-related factors. Each signal has documented predictive power in the academic defect prediction literature.
| # | Signal | Category | What It Measures | Citation |
|---|---|---|---|---|
| 01 | Database Migration Presence (DDL) | Schema & Infrastructure | ALTER TABLE, DROP COLUMN, or index creation in the diff | Kim et al., MSR 2008 |
| 02 | Historical Failure Rate | Historical Patterns | Repository or service failure rate above historical baseline | Kim & Whitehead, MSR 2008 |
| 03 | Change Entropy | Change Complexity | Shannon entropy measuring dispersion of changes across subsystems | Hassan, ICSE 2009 |
| 04 | Author File Expertise | Author & Review | Prior experience with the specific files being changed | Bird et al., FSE 2011 |
| 05 | SLO Error Budget Burn Rate | Deployment Context | Error budget consumption rate at deploy time | Beyer et al., Google SRE Book 2016 |
| 06 | Minor Contributors | Author & Review | Density of first-time file contributors among authors | Bird et al., FSE 2011 |
| 07 | LLM Semantic Risk Analysis | AI & Emerging | Risk phrases in PR title, description, or review comments | Koalr internal |
| 08 | Review Coverage | Author & Review | Reviewer count and cross-team approval status | Bacchelli & Bird, ICSE 2013 |
| 09 | Change Burst | Change Complexity | Rapid successive commits to the same file | Nagappan & Zeller, ICSE 2010 |
| 10 | File Count & Directory Spread | Change Complexity | Number of files changed and directory breadth | Hassan & Holt, ICSM 2005 |
| 11 | Patch Coverage Delta | Test & Coverage | Test coverage change on modified lines | Inozemtseva & Holmes, ICSE 2014 |
| 12 | Code Churn Rate | Change Complexity | File modification frequency over a trailing window | Nagappan & Ball, ICSE 2005 |
| 13 | Blast Radius (Downstream Service Count) | Deployment Context | Number of downstream services affected | Beyer et al., Google SRE Book 2016 |
| 14 | Deployment Frequency Inverse | Historical Patterns | Time elapsed since last deploy to this service | Forsgren et al., Accelerate 2018 |
| 15 | CI / Test Flakiness Rate | Test & Coverage | Flaky test rate for the affected service | Lam et al., ICSM 2019 |
| 16 | Static Analysis Violation Delta | Test & Coverage | New security hotspots or critical violations introduced | Palomba et al., TSE 2018 |
| 17 | Dependency Major Version Bumps | Schema & Infrastructure | Major semver bumps in direct dependencies | Bavota et al., JSS 2015 |
| 18 | Team Historical CFR | Historical Patterns | Team change failure rate over trailing window | Forsgren et al., Accelerate 2018 |
| 19 | Recent Incidents in Service | Historical Patterns | Recent incidents attributed to this service | Kim & Whitehead, MSR 2008 |
| 20 | PR Age (Time Open Before Merge) | Author & Review | Very short (insufficient review) or very long (context drift) PR age | Rigby & Bird, FSE 2013 |
| 21 | CVE Risk in New Dependencies | Schema & Infrastructure | Newly introduced dependency with known vulnerabilities | Decan et al., MSR 2018 |
| 22 | Rollback Frequency | Historical Patterns | Service rollback rate over trailing window | Kim et al., MSR 2008 |
| 23 | Number of Reviewers | Author & Review | Reviewer count relative to change size | Rigby & Bird, FSE 2013 |
| 24 | Lines Added (Change Size) | Change Complexity | Volume of additions in non-generated files | Nagappan & Ball, ICSE 2005 |
| 25 | Deployment Timing (Day / Hour) | Deployment Context | Late-week or after-hours deployment timing | Forsgren et al., Accelerate 2018 |
| 26 | Review Depth (Comment Count) | Author & Review | Review comment count relative to change size | Bacchelli & Bird, ICSE 2013 |
| 27 | Cyclomatic / Cognitive Complexity Delta | Change Complexity | Complexity increase across changed files | Eaddy et al., ICSE 2008 |
| 28 | AI Authorship Ratio | AI & Emerging | Proportion of AI-generated code without review iteration | Koalr internal |
| 29 | Developer Graph Centrality | AI & Emerging | Author position in co-commit collaboration graph | Xia et al., PMC 2023 |
| 30 | Author Service Commit Recency | Author & Review | Time since last author commit to this service | Mockus & Herbsleb, ICSE 2002 |
| 31 | Lines Deleted | Change Complexity | Volume of deletions without corresponding test additions | Nagappan & Ball, ICSE 2005 |
| 32 | Collaboration Graph Centrality | Author & Review | Authors who are structural bridges between teams (high betweenness centrality) carry higher deploy risk | Xia et al., PMC 2023 |
| 33 | Concurrent PR Merge Risk | Deployment Context | Other PRs merging to the same repo within ±2 hours multiply blast surface | Kochhar et al., MSR 2015 |
| 34 | IaC / Pipeline Change Detection | Schema & Infrastructure | Terraform, Kubernetes manifests, Helm charts, Dockerfile, or GitHub Actions changes (hard-gated at 75) | Rahman & Williams, MSR 2019 |
| 35 | Tangled Commit Detection | Change Complexity | PRs mixing unrelated concerns (feature + refactor + bug fix) have 2× higher post-release defect density than focused changes | Herzig et al., MSR 2013 |
| 36 | Reviewer File Familiarity | Author & Review | Reviewers with prior commit history in the changed files catch significantly more defects than reviewers unfamiliar with the code | Rigby & Bird 2013; McIntosh et al., EMSE 2015 |
Diffusion matters as much as size. Change entropy captures something raw line counts miss. A 100-line change across 8 services is higher risk than a 1,000-line change in a single service.
Coverage delta trumps overall coverage. Aggregate test coverage percentage has limited correlation with incident probability. What matters is the coverage change on modified lines specifically.
Signals interact non-linearly. A PR with multiple moderate signals active is higher risk than any single signal suggests. This is why ML-based scoring outperforms simple threshold rules.
How signals combine into a 0–100 score
Proceed with normal process. Small, well-reviewed changes by experienced authors with no active SLO pressure.
Worth a second look. An additional reviewer, coverage improvement, or timing adjustment often drops the score below 40.
Strong signals present. Consider decomposing the PR, adding reviewers, writing missing tests, or deferring to a lower-risk deploy window.
See these signals score your next deploy
Connect GitHub in under 5 minutes. Koalr calculates a risk score on every open PR immediately — multiple research-validated signals combined into a single 0–100 score with per-signal breakdown and suggested actions.
Related reading
Deployment Risk vs. Delivery Risk
Why these are different questions — and why delivery metrics cannot predict individual deploy failures.
Deployment Risk Management: A Practical Guide
Feature flags, canary releases, rollbacks, and how automated risk scoring changes the equation.
Change Entropy Explained
The Shannon entropy signal — why diffusion beats size as a risk predictor.
Author Expertise Signals
How Koalr scores per-file expertise and why experience in the right files matters more than seniority.