Ship production releases with confidence and safety
A pre-launch checklist plus feature-flagged staged rollout with metric-based advance/hold/rollback thresholds and a rollback plan template.
Why it matters
Deploy features to production safely with comprehensive pre-launch verification, staged rollouts, monitoring, and rollback plans to ensure every release is reversible, observable, and incremental.
Outcomes
What it gets done
Verify code quality, security, performance, and accessibility against comprehensive checklists before deployment
Deploy behind feature flags with gradual rollout from 5% to 100% of users while monitoring thresholds
Monitor error rates, latency, and business metrics at each rollout stage with clear advance/hold/rollback criteria
Execute rollback plans immediately when error rates exceed 2x baseline or latency degrades beyond thresholds
Install
Add it to your toolbox
Run in your project directory:
curl -fsSL https://spark.entire.vc/get/ag-shipping-and-launch | bash Overview
Shipping and Launch
This skill runs a shipping process: a 6-category pre-launch checklist, feature-flagged staged rollout (team, canary 5%, 25/50/100%) with metric-based advance/hold/rollback thresholds, layered monitoring, and a written rollback plan with stated rollback times per mechanism. Use it for a first production deployment, a significant user-facing release, a data/infrastructure migration, a beta launch, or any deployment carrying risk.
What it does
A shipping and launch process built on making every deploy reversible, observable, and incremental. A pre-launch checklist covers six categories: code quality (all tests pass, clean build/lint, code reviewed, no debug console.log or unresolved TODOs); security (no committed secrets, no critical/high npm audit findings, input validation, auth checks, security headers, rate limiting on auth endpoints, CORS scoped to specific origins); performance (Core Web Vitals within Good thresholds, no N+1 queries, optimized images, bundle size budget, indexed queries, caching); accessibility (keyboard navigation, screen reader support, WCAG 2.1 AA color contrast, correct focus management, descriptive error messages, no axe-core/Lighthouse warnings); infrastructure (env vars set, migrations applied, DNS/SSL/CDN configured, logging and error reporting wired up, a working health check endpoint); and documentation (README, API docs, ADRs, changelog, user-facing docs all current). Features ship behind flags to decouple deployment from release, following a five-stage lifecycle - deploy with the flag off, enable for team/beta, gradual rollout (5% to 25% to 50% to 100%), monitor at every stage, then clean up the flag and dead code path - with rules that every flag has an owner and expiration date, cleanup happens within 2 weeks of full rollout, flags are never nested (exponential state combinations), and both flag states are tested in CI. The staged rollout sequence runs staging deploy with full test suite, production deploy with the flag off (health check plus error-monitoring verification), team-only enablement with a 24-hour monitoring window, a canary at 5% of users with a 24-48 hour window comparing metrics against baseline, gradual increase through 25/50/100% with the same monitoring and rollback-to-previous-percentage ability at each step, and finally full rollout with a week of monitoring before flag cleanup. A four-metric threshold table decides advance/hold/rollback at each stage: error rate (green within 10% of baseline, yellow 10-100% above, red over 2x baseline), P95 latency (green within 20%, yellow 20-50% above, red over 50% above), client JS errors (green no new types, yellow new errors under 0.1% of sessions, red over 0.1%), and business metrics (green neutral/positive, yellow under 5% decline as possible noise, red over 5% decline) - with immediate rollback triggers for error rate beyond 2x baseline, P95 latency beyond 50% above baseline, a spike in user-reported issues, data integrity problems, or a discovered security vulnerability. Monitoring spans application metrics (error rate, response time percentiles, request volume, active users, business metrics), infrastructure metrics (CPU/memory, DB connection pool, disk, network latency, queue depth), and client metrics (Core Web Vitals, JS errors, client-perceived API error rate, page load time), with error reporting wired through both a React error boundary and server-side middleware that never exposes internals to users. Post-launch verification in the first hour checks the health endpoint, error and latency dashboards, manually tests the critical user flow, confirms logs are flowing, and verifies the rollback mechanism actually works. Every deployment gets a written rollback plan before shipping - trigger conditions, rollback steps (flag disable or git revert and redeploy), database migration rollback considerations, and a stated time-to-rollback per mechanism (flag under 1 minute, redeploy under 5, database under 15).
When to use - and when NOT to
Use it for a first production deployment of a feature, releasing a significant change to users, migrating data or infrastructure, opening a beta program, or any deployment carrying risk - which the skill treats as all of them.
Inputs and outputs
Input is a feature or change ready to ship. Output is a completed pre-launch checklist, a feature flag with a defined rollout lifecycle, a staged rollout plan with metric thresholds per stage, configured monitoring across application/infrastructure/client layers, and a written rollback plan with stated rollback times.
Integrations
References a project-wide Definition of Done plus separate security, performance, and accessibility pre-launch checklists; error reporting wires through a React error boundary and server middleware to an error tracking service.
Who it's for
Developers and teams shipping a risky or significant change who want a systematic pre-launch checklist, a feature-flagged staged rollout with objective advance/hold/rollback thresholds, and a rollback plan written before anything ships rather than improvised after something breaks.
Source README
Shipping and Launch
Overview
Ship with confidence. The goal is not just to deploy - it's to deploy safely, with monitoring in place, a rollback plan ready, and a clear understanding of what success looks like. Every launch should be reversible, observable, and incremental.
When to Use
- Deploying a feature to production for the first time
- Releasing a significant change to users
- Migrating data or infrastructure
- Opening a beta or early access program
- Any deployment that carries risk (all of them)
The Pre-Launch Checklist
Code Quality
- All tests pass (unit, integration, e2e)
- Build succeeds with no warnings
- Lint and type checking pass
- Code reviewed and approved
- No TODO comments that should be resolved before launch
- No
console.logdebugging statements in production code - Error handling covers expected failure modes
Security
- No secrets in code or version control
-
npm auditshows no critical or high vulnerabilities - Input validation on all user-facing endpoints
- Authentication and authorization checks in place
- Security headers configured (CSP, HSTS, etc.)
- Rate limiting on authentication endpoints
- CORS configured to specific origins (not wildcard)
Performance
- Core Web Vitals within "Good" thresholds
- No N+1 queries in critical paths
- Images optimized (compression, responsive sizes, lazy loading)
- Bundle size within budget
- Database queries have appropriate indexes
- Caching configured for static assets and repeated queries
Accessibility
- Keyboard navigation works for all interactive elements
- Screen reader can convey page content and structure
- Color contrast meets WCAG 2.1 AA (4.5:1 for text)
- Focus management correct for modals and dynamic content
- Error messages are descriptive and associated with form fields
- No accessibility warnings in axe-core or Lighthouse
Infrastructure
- Environment variables set in production
- Database migrations applied (or ready to apply)
- DNS and SSL configured
- CDN configured for static assets
- Logging and error reporting configured
- Health check endpoint exists and responds
Documentation
- README updated with any new setup requirements
- API documentation current
- ADRs written for any architectural decisions
- Changelog updated
- User-facing documentation updated (if applicable)
Feature Flag Strategy
Ship behind feature flags to decouple deployment from release:
// Feature flag check
const flags = await getFeatureFlags(userId);
if (flags.taskSharing) {
// New feature: task sharing
return <TaskSharingPanel task={task} />;
}
// Default: existing behavior
return null;
Feature flag lifecycle:
1. DEPLOY with flag OFF → Code is in production but inactive
2. ENABLE for team/beta → Internal testing in production environment
3. GRADUAL ROLLOUT → 5% → 25% → 50% → 100% of users
4. MONITOR at each stage → Watch error rates, performance, user feedback
5. CLEAN UP → Remove flag and dead code path after full rollout
Rules:
- Every feature flag has an owner and an expiration date
- Clean up flags within 2 weeks of full rollout
- Don't nest feature flags (creates exponential combinations)
- Test both flag states (on and off) in CI
Staged Rollout
The Rollout Sequence
1. DEPLOY to staging
└── Full test suite in staging environment
└── Manual smoke test of critical flows
2. DEPLOY to production (feature flag OFF)
└── Verify deployment succeeded (health check)
└── Check error monitoring (no new errors)
3. ENABLE for team (flag ON for internal users)
└── Team uses the feature in production
└── 24-hour monitoring window
4. CANARY rollout (flag ON for 5% of users)
└── Monitor error rates, latency, user behavior
└── Compare metrics: canary vs. baseline
└── 24-48 hour monitoring window
└── Advance only if all thresholds pass (see table below)
5. GRADUAL increase (25% -> 50% -> 100%)
└── Same monitoring at each step
└── Ability to roll back to previous percentage at any point
6. FULL rollout (flag ON for all users)
└── Monitor for 1 week
└── Clean up feature flag
Rollout Decision Thresholds
Use these thresholds to decide whether to advance, hold, or roll back at each stage:
| Metric | Advance (green) | Hold and investigate (yellow) | Roll back (red) |
|---|---|---|---|
| Error rate | Within 10% of baseline | 10-100% above baseline | >2x baseline |
| P95 latency | Within 20% of baseline | 20-50% above baseline | >50% above baseline |
| Client JS errors | No new error types | New errors at <0.1% of sessions | New errors at >0.1% of sessions |
| Business metrics | Neutral or positive | Decline <5% (may be noise) | Decline >5% |
When to Roll Back
Roll back immediately if:
- Error rate increases by more than 2x baseline
- P95 latency increases by more than 50%
- User-reported issues spike
- Data integrity issues detected
- Security vulnerability discovered
Monitoring and Observability
What to Monitor
Application metrics:
├── Error rate (total and by endpoint)
├── Response time (p50, p95, p99)
├── Request volume
├── Active users
└── Key business metrics (conversion, engagement)
Infrastructure metrics:
├── CPU and memory utilization
├── Database connection pool usage
├── Disk space
├── Network latency
└── Queue depth (if applicable)
Client metrics:
├── Core Web Vitals (LCP, INP, CLS)
├── JavaScript errors
├── API error rates from client perspective
└── Page load time
Error Reporting
// Set up error boundary with reporting
class ErrorBoundary extends React.Component {
componentDidCatch(error: Error, info: React.ErrorInfo) {
// Report to error tracking service
reportError(error, {
componentStack: info.componentStack,
userId: getCurrentUser()?.id,
page: window.location.pathname,
});
}
render() {
if (this.state.hasError) {
return <ErrorFallback onRetry={() => this.setState({ hasError: false })} />;
}
return this.props.children;
}
}
// Server-side error reporting
app.use((err: Error, req: Request, res: Response, next: NextFunction) => {
reportError(err, {
method: req.method,
url: req.url,
userId: req.user?.id,
});
// Don't expose internals to users
res.status(500).json({
error: { code: 'INTERNAL_ERROR', message: 'Something went wrong' },
});
});
Post-Launch Verification
In the first hour after launch:
1. Check health endpoint returns 200
2. Check error monitoring dashboard (no new error types)
3. Check latency dashboard (no regression)
4. Test the critical user flow manually
5. Verify logs are flowing and readable
6. Confirm rollback mechanism works (dry run if possible)
Rollback Strategy
Every deployment needs a rollback plan before it happens:
### Rollback Plan for [Feature/Release]
### Trigger Conditions
- Error rate > 2x baseline
- P95 latency > [X]ms
- User reports of [specific issue]
### Rollback Steps
1. Disable feature flag (if applicable)
OR
1. Deploy previous version: `git revert <commit> && git push`
2. Verify rollback: health check, error monitoring
3. Communicate: notify team of rollback
### Database Considerations
- Migration [X] has a rollback: `npx prisma migrate rollback`
- Data inserted by new feature: [preserved / cleaned up]
### Time to Rollback
- Feature flag: < 1 minute
- Redeploy previous version: < 5 minutes
- Database rollback: < 15 minutes
See Also
- For the project-wide Definition of Done that every change must clear before this checklist, see
references/definition-of-done.md - For security pre-launch checks, see
references/security-checklist.md - For performance pre-launch checklist, see
references/performance-checklist.md - For accessibility verification before launch, see
references/accessibility-checklist.md
Common Rationalizations
| Rationalization | Reality |
|---|---|
| "It works in staging, it'll work in production" | Production has different data, traffic patterns, and edge cases. Monitor after deploy. |
| "We don't need feature flags for this" | Every feature benefits from a kill switch. Even "simple" changes can break things. |
| "Monitoring is overhead" | Not having monitoring means you discover problems from user complaints instead of dashboards. |
| "We'll add monitoring later" | Add it before launch. You can't debug what you can't see. |
| "Rolling back is admitting failure" | Rolling back is responsible engineering. Shipping a broken feature is the failure. |
Red Flags
- Deploying without a rollback plan
- No monitoring or error reporting in production
- Big-bang releases (everything at once, no staging)
- Feature flags with no expiration or owner
- No one monitoring the deploy for the first hour
- Production environment configuration done by memory, not code
- "It's Friday afternoon, let's ship it"
Verification
Before deploying:
- Pre-launch checklist completed (all sections green)
- Feature flag configured (if applicable)
- Rollback plan documented
- Monitoring dashboards set up
- Team notified of deployment
After deploying:
- Health check returns 200
- Error rate is normal
- Latency is normal
- Critical user flow works
- Logs are flowing
- Rollback tested or verified ready
Limitations
- Use this skill only when the task clearly matches its upstream source and local project context.
- Verify commands, generated code, dependencies, credentials, and external service behavior before applying changes.
- Do not treat examples as a substitute for environment-specific tests, security review, or user approval for destructive or costly actions.
FAQ
Common questions
Discussion
Questions & comments · 0
Sign In Sign in to leave a comment.