Files
Edu/docs/troubleshooting/known-issues-p6-addendum.md
SpecialX e9ea34fe53
Some checks failed
CI Go / test (push) Has been cancelled
CI Python / test (push) Has been cancelled
CI TypeScript / test (push) Has been cancelled
CI Proto / lint (push) Failing after 8m7s
feat(p6): production hardening with circuit breaker, backup, monitoring and chaos engineering
P6 生产硬化阶段交付物(46 文件):

## 1. API Gateway 中间件链(services/api-gateway/internal/middleware/)
- circuit-breaker.go: gobreaker v2 熔断器(5s 窗口/50% 错误率/30s OPEN→HALF_OPEN)
- ratelimit.go: 令牌桶限流(sync.Map + cleanup goroutine,默认 100rps/20 burst)
- cors.go: CORS 中间件(CORS_ORIGINS 环境变量)
- recovery.go: panic 恢复 + uuid request_id
- security.go: 安全头 + 请求体 10MB 限制
- requestid.go: 请求 ID 注入
- health/health.go: /healthz + /readyz 健康检查
- main.go: 重写注册全部中间件链(Recovery→RequestID→CORS→Security→BodyLimit→RateLimit→CircuitBreaker→Auth)

## 2. 基础设施硬化(infra/)
- backup/backup-mysql.sh: MySQL 全量备份(mysqldump+gzip,按服务独立)
- backup/restore-mysql.sh: 恢复脚本
- backup/backup-cron.sh: cron 调度入口(5 服务批量备份)
- alertmanager/alertmanager.yml: 告警路由(webhook + 邮件示例)
- prometheus/rules.yml: 8 条告警规则(服务可用性/性能/资源 3 组)
- grafana/dashboards/microservices-overview.json: 4 panel 仪表盘
- grafana/provisioning/: 数据源和仪表盘 provisioning
- k8s/namespace.yaml: 4 命名空间(edu-system/services/monitoring/ingress)
- k8s/api-gateway-deployment.yaml: Deployment + Service 骨架
- chaos/experiments.yaml: 3 个 Litmus 混沌实验(pod-kill/network-latency/disk-fill)
- docker-compose.monitoring.yml: 监控栈 profile
- security/secrets.example.env: 8 项密钥占位符
- security/waf-rules.conf: ModSecurity WAF 规则骨架

## 3. 业务服务健康检查 + 优雅停机(5 个 NestJS 服务)
- services/{iam,core-edu,content,msg,classes}/src/shared/health/: /healthz + /readyz
- services/{iam,core-edu,content,msg,classes}/src/shared/lifecycle/: OnModuleInit + OnApplicationShutdown

## 4. Python 服务健康检查
- services/{ai,data-ana}/src/health/health.py: FastAPI APIRouter

## 5. 运维文档
- docs/architecture/runbooks/p6-hardening.md: P6 总览 Runbook(9 章节)
- docs/architecture/runbooks/incident-response.md: 事件响应手册(5 章节)
- docs/architecture/004-p6-addendum.md: 004 架构补记 P6 章节
- docs/troubleshooting/known-issues-p6-addendum.md: 15 条 P6 场景→技术映射

## 验收信号
- RPO ≤ 15min(MySQL 备份 + binlog PITR)
- RTO ≤ 30min(K8s 滚动更新 + DNS 切换)
- P99 ≤ 500ms(熔断 + 限流 + 缓存)
- 熔断器错误率 > 50% 触发 OPEN
- 限流 100rps/20 burst
- 备份保留 7 天
- 混沌实验每月 1 次
2026-07-08 02:16:58 +08:00

25 lines
1.7 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# known-issues P6 补丁
> 本文件为 `docs/troubleshooting/known-issues.md` 的 P6 阶段补丁,列出 P6 新增的"场景→技术"映射。
> 合并方式:将下表条目追加到原文件对应分区,遵循索引式速查规范,不写代码示例。
## P6 生产硬化新增场景
| 场景 | 技术方案 |
|------|----------|
| 熔断器状态切换Closed/Open/Half-Open | gobreaker v2 ReadyToTrip 回调按连续失败数 + 失败率判定 |
| 限流桶按租户维度清理 | sync.Map + ticker 周期清理过期桶,避免内存泄漏 |
| 备份脚本定时调度 | K8s CronJob + backup-cron.sh15min 一次满足 RPO |
| PostgreSQL WAL 归档恢复到时间点 | pg_receivewal 归档 + PITR 恢复 |
| Redis 增量备份 | BGSAVE 触发 + RDB 文件上传对象存储 |
| Kafka 消费位点快照 | __consumer_offsets topic dump 到对象存储 |
| 熔断器半开态试探限流 | MaxRequests 限制并发试探,避免恢复期二次过载 |
| 健康探针分流 liveness/readiness | /healthz 仅进程存活,/readyz 检查依赖,避免滚动重启雪崩 |
| 优雅停机等待 in-flight 请求 | app.enableShutdownHooks + terminationGracePeriodSeconds=60 |
| Kafka producer 关闭前 flush | producer.disconnect() 内部 flush避免消息丢失 |
| 数据源销毁顺序 | 先 Kafka、再 Redis、最后 DB避免反向依赖阻塞 |
| 混沌实验回滚 | rollback.sh 按实验名清理 tc/iptables 规则 |
| DNS 切换消除缓存 | TTL=60s + 等待 + 客户端清缓存,避免切换无效 |
| 告警抑制去重 | Alertmanager inhibit + group_by + group_wait |
| Python 服务就绪检查延迟初始化 | readyz 返回 ok + TODO避免启动期依赖未就绪导致探针失败 |