Recovery Automation vs. Failure Propagation Risk
Bound automated recovery to staged, blast-radius-limited domains with mandatory rollback and independent health validation before expanding remediation scope.
CyberTRIZ analysis · Telecommunications contradiction RO031 · one of 8,235 worked contradictions published by CyberTRIZ.AI
Regulations
Business Context
Automated recovery can detect failures, reroute traffic, restart functions, scale resources, modify configurations, and restore services much faster than manual operations. However, an incorrect diagnosis or recovery rule can execute across many network elements before human operators recognize the problem. Automation therefore reduces restoration time while potentially increasing the speed and scale of failure propagation.
Telecommunications TRIZ Resolution
Recovery automation should operate within bounded domains and progressively expand only when results confirm that the action is effective. Canary recovery, staged execution, predefined blast-radius limits, automated rollback, independent health checks, and confidence thresholds can prevent a single incorrect action from affecting the complete network.
Applicable TRIZ Principles
Principle 1 – Segmentation limits automated recovery actions to controlled failure domains.
Principle 11 – Beforehand Cushioning establishes rollback and containment mechanisms before automation executes.
Principle 23 – Feedback verifies each recovery stage before expanding the action.
Expected Outcome
Faster automated recovery
Smaller automation failure impact
Greater confidence in closed-loop operations
Reduced risk of network-wide propagation
Decision Indicators
Early indicators include:
Automated recovery actions can modify large network domains simultaneously.
Recovery logic lacks independent outcome validation.
Operators disable automation after high-impact incidents.
Rollback depends on manual intervention.
A single incorrect diagnosis can trigger repeated automated actions across the network.