CMDB 资产管理(3):云资源发现、同步与变更历史
这是“CMDB 资产管理”连续系列的第 3 篇,也是完整实现教程的第 2 篇,只覆盖第 10~18 章。上一篇已经完成项目骨架、账号体系、CMDB 基础资产模型与基础页面;本篇把“云厂商返回的数据”安全地变成“可审计、可重复执行、能表达缺失与恢复”的 CMDB 状态。
本篇所有文件内容均来自最终实现。读者侧项目根目录统一写作 devopsX/。命令按一次一条执行;每条命令前都给出执行目录。文中不要求真实云凭据,Fake Provider 足以完成全部核心闭环。
10 同步审计模型:SyncRun 与 ComputeInstanceChange
10.1 先把同步术语讲清楚
先定义名词,再进入锁、事务与生命周期。下面这些词会贯穿第 10~18 章。
| 术语 | 本项目中的准确含义 | 不代表什么 |
|---|---|---|
| 适配器(adapter) | 隔离不同云厂商 SDK 差异的对象,对服务层只暴露 discover(account)。 | 不是 Django Model,也不直接写数据库。 |
| DTO | Data Transfer Object,数据传输对象。本项目用冻结的 dataclass 表达地域、可用区、实例与标签。 | 不是 ORM 对象,不带 save()。 |
| 快照(snapshot) | 某次发现返回的一组 DTO,以及完整性、错误与发现范围元数据。 | 不一定代表全账号全部资源;要结合 scope 与 complete 判断。 |
| 范围(scope) | 快照声称自己覆盖“全部地域”还是“账号白名单中的地域”。 | 不是页面筛选条件。 |
| 部分结果(partial result) | 已经取得一部分合法数据,但发现过程没有完整结束的快照,complete=False。 | 不能用“未返回”推断资源已消失。 |
| 对账(reconciliation) | 把快照与数据库当前状态比较,创建、更新、标记缺失或恢复,并留下变化记录。 | 不是先清空再全量插入。 |
| 幂等性(idempotence) | 同一完整快照重复执行,不重复创建资产,也不重复生成业务变化记录。 | 不表示完全没有数据库写入;最后发现时间仍会刷新。 |
| 生命周期(lifecycle) | 实例在 CMDB 中的 present、missing、retired 状态流转。 | 不同于云厂商的 Running、Stopped 等运行状态。 |
| 事务(transaction) | 一组数据库操作要么一起提交,要么一起回滚。 | 不能把外部网络调用也变成数据库原子操作。 |
| 行锁(row lock) | MySQL 中 SELECT FOR UPDATE 对选中行施加的事务锁。 | 不是分布式锁;SQLite 也不提供同等语义。 |
| 过期执行(stale run) | 等待或运行超过 30 分钟仍未结束的 SyncRun,视为遗留阻塞。 | 不等于云资源过期。 |
| 栅栏校验(fencing) | 写入快照前再次确认“本次 SyncRun 仍为 running”且同步配置没有变化;旧执行或基于旧配置取得的快照都禁止落库。 | 不是令牌租约系统,也不是跨服务分布式一致性协议。 |
一条完整链路可以概括为:适配器把厂商响应转成 DTO,组成带 scope 的 snapshot;服务层先验证快照,再在事务中 reconciliation;SyncRun 记录一次执行,ComputeInstanceChange 记录每台实例的业务变化。
10.2 两类审计记录分别回答什么问题
10.2.1 SyncRun 记录一次执行
public_id是对外展示的 UUID,页面不必暴露连续数据库主键。status区分等待、运行、成功、部分成功和失败。trigger区分页面、管理命令和 API;同一服务函数因此可以被不同入口复用。- 七个统计字段分别记录发现、新建、更新、未变化、缺失、恢复和退役数量。
error_code用于稳定判断,error_message用于显示脱敏后的说明。
10.2.2 ComputeInstanceChange 记录一台实例发生了什么
action只保存创建、更新、标记缺失、恢复和退役五种业务动作。changed_fields保存发生变化的字段名。before_data与after_data保存可审计状态,而不是依赖以后再反推。sync_run允许为空,因为人工退役可以不依附某次云发现。- 外键使用
PROTECT,避免删除实例或执行记录时把审计链一起抹掉。
10.3 最终模型文件
为了保证模型、约束、权限和审计字段可以整体复制,下面给出最终完整文件。上一篇已经讲过的基础模型也保留在文件中;本章重点阅读 ComputeInstanceTag、SyncRun 和 ComputeInstanceChange。
相对路径:devopsX/cmdb/models.py;内容:完整文件。
import re
import uuid
from django.conf import settings
from django.core.exceptions import ValidationError
from django.db import models
from django.utils import timezone
def validate_credential_profile(value):
if value and not re.fullmatch(r"[A-Z][A-Z0-9_]*", value):
raise ValidationError("凭据前缀只能使用大写字母、数字和下划线。")
def validate_region_allowlist(value):
if not isinstance(value, list):
raise ValidationError("地域白名单必须是地域 ID 列表。")
normalized = []
for item in value:
if not isinstance(item, str) or not item.strip():
raise ValidationError("地域白名单只能包含非空字符串。")
region_id = item.strip()
if len(region_id) > 100:
raise ValidationError("地域 ID 不能超过 100 个字符。")
if region_id in normalized:
raise ValidationError("地域白名单不能包含重复地域 ID。")
normalized.append(region_id)
if normalized != value:
raise ValidationError("地域 ID 首尾不能包含空格。")
class CloudProvider(models.Model):
code = models.SlugField("代码", max_length=32, unique=True)
name = models.CharField("名称", max_length=100)
is_active = models.BooleanField("启用", default=True)
created_at = models.DateTimeField("创建时间", auto_now_add=True)
updated_at = models.DateTimeField("更新时间", auto_now=True)
class Meta:
ordering = ["code"]
verbose_name = "云厂商"
verbose_name_plural = "云厂商"
def __str__(self):
return "%s - %s" % (self.code, self.name)
class CloudAccount(models.Model):
provider = models.ForeignKey(
CloudProvider,
on_delete=models.PROTECT,
related_name="accounts",
verbose_name="云厂商",
)
account_key = models.SlugField("账号键", max_length=100)
name = models.CharField("显示名称", max_length=100)
credential_profile = models.CharField(
"凭据环境变量前缀",
max_length=100,
blank=True,
validators=[validate_credential_profile],
)
region_allowlist = models.JSONField(
"地域白名单",
default=list,
blank=True,
validators=[validate_region_allowlist],
)
is_active = models.BooleanField("启用", default=True)
sync_enabled = models.BooleanField("允许同步", default=True)
last_successful_sync_at = models.DateTimeField(
"最后成功同步时间",
null=True,
blank=True,
)
created_at = models.DateTimeField("创建时间", auto_now_add=True)
updated_at = models.DateTimeField("更新时间", auto_now=True)
class Meta:
ordering = ["provider__code", "account_key"]
constraints = [
models.UniqueConstraint(
fields=["provider", "account_key"],
name="cmdb_unique_provider_account_key",
)
]
permissions = [
("sync_cloudaccount", "可以同步云账号"),
("import_cloudaccount", "可以导入云账号"),
]
verbose_name = "云账号"
verbose_name_plural = "云账号"
def __str__(self):
return "%s - %s" % (self.provider.code, self.name)
class CloudRegion(models.Model):
account = models.ForeignKey(
CloudAccount,
on_delete=models.PROTECT,
related_name="regions",
verbose_name="云账号",
)
provider_resource_id = models.CharField("厂商地域 ID", max_length=100)
name = models.CharField("名称", max_length=100)
endpoint = models.CharField("服务端点", max_length=255, blank=True)
is_active = models.BooleanField("当前存在", default=True)
first_seen_at = models.DateTimeField("首次发现时间", default=timezone.now)
last_seen_at = models.DateTimeField("最后发现时间", default=timezone.now)
class Meta:
ordering = ["account", "provider_resource_id"]
constraints = [
models.UniqueConstraint(
fields=["account", "provider_resource_id"],
name="cmdb_unique_account_region_id",
)
]
verbose_name = "云地域"
verbose_name_plural = "云地域"
def __str__(self):
return "%s - %s" % (self.account.account_key, self.name)
class CloudAvailabilityZone(models.Model):
account = models.ForeignKey(
CloudAccount,
on_delete=models.PROTECT,
related_name="availability_zones",
verbose_name="云账号",
)
region = models.ForeignKey(
CloudRegion,
on_delete=models.PROTECT,
related_name="availability_zones",
verbose_name="云地域",
)
provider_resource_id = models.CharField("厂商可用区 ID", max_length=100)
name = models.CharField("名称", max_length=100)
is_active = models.BooleanField("当前存在", default=True)
first_seen_at = models.DateTimeField("首次发现时间", default=timezone.now)
last_seen_at = models.DateTimeField("最后发现时间", default=timezone.now)
class Meta:
ordering = ["account", "region", "provider_resource_id"]
constraints = [
models.UniqueConstraint(
fields=["account", "provider_resource_id"],
name="cmdb_unique_account_zone_id",
)
]
verbose_name = "可用区"
verbose_name_plural = "可用区"
def clean(self):
super().clean()
if self.region_id and self.account_id:
try:
region = self.region
except CloudRegion.DoesNotExist:
region = None
if region is not None and region.account_id != self.account_id:
raise ValidationError("可用区账号必须与地域账号一致。")
def __str__(self):
return "%s - %s" % (self.region.provider_resource_id, self.name)
class ComputeInstance(models.Model):
class NormalizedStatus(models.TextChoices):
RUNNING = "running", "运行中"
STOPPED = "stopped", "已停止"
STARTING = "starting", "启动中"
STOPPING = "stopping", "停止中"
UNKNOWN = "unknown", "未知"
class LifecycleState(models.TextChoices):
PRESENT = "present", "当前存在"
MISSING = "missing", "本次未发现"
RETIRED = "retired", "已退役"
account = models.ForeignKey(
CloudAccount,
on_delete=models.PROTECT,
related_name="compute_instances",
verbose_name="云账号",
)
region = models.ForeignKey(
CloudRegion,
on_delete=models.PROTECT,
related_name="compute_instances",
verbose_name="云地域",
)
availability_zone = models.ForeignKey(
CloudAvailabilityZone,
on_delete=models.PROTECT,
related_name="compute_instances",
verbose_name="可用区",
null=True,
blank=True,
)
provider_resource_id = models.CharField("厂商实例 ID", max_length=100)
name = models.CharField("实例名称", max_length=255, blank=True)
instance_type = models.CharField("实例规格", max_length=100, blank=True)
vcpu = models.PositiveIntegerField("vCPU", default=0)
memory_mb = models.PositiveIntegerField("内存 MiB", default=0)
os_name = models.CharField("操作系统", max_length=255, blank=True)
provider_status = models.CharField("厂商原始状态", max_length=100, blank=True)
normalized_status = models.CharField(
"标准状态",
max_length=20,
choices=NormalizedStatus.choices,
default=NormalizedStatus.UNKNOWN,
)
lifecycle_state = models.CharField(
"生命周期",
max_length=20,
choices=LifecycleState.choices,
default=LifecycleState.PRESENT,
)
private_ips = models.JSONField("私网 IP", default=list, blank=True)
public_ips = models.JSONField("公网 IP", default=list, blank=True)
cloud_created_at = models.DateTimeField("云上创建时间", null=True, blank=True)
first_seen_at = models.DateTimeField("首次发现时间", default=timezone.now)
last_seen_at = models.DateTimeField("最后发现时间", default=timezone.now)
missing_since = models.DateTimeField("开始缺失时间", null=True, blank=True)
retired_at = models.DateTimeField("退役时间", null=True, blank=True)
created_at = models.DateTimeField("创建时间", auto_now_add=True)
updated_at = models.DateTimeField("更新时间", auto_now=True)
class Meta:
ordering = ["account", "provider_resource_id"]
constraints = [
models.UniqueConstraint(
fields=["account", "provider_resource_id"],
name="cmdb_unique_account_instance_id",
)
]
indexes = [
models.Index(
fields=["lifecycle_state", "normalized_status"],
name="cmdb_inst_life_status_idx",
),
models.Index(fields=["name"], name="cmdb_inst_name_idx"),
]
permissions = [
("export_computeinstance", "可以导出计算实例"),
("retire_computeinstance", "可以退役计算实例"),
]
verbose_name = "计算实例"
verbose_name_plural = "计算实例"
def clean(self):
super().clean()
region = None
if self.region_id and self.account_id:
try:
region = self.region
except CloudRegion.DoesNotExist:
pass
if region is not None and region.account_id != self.account_id:
raise ValidationError("实例账号必须与地域账号一致。")
if self.availability_zone_id:
try:
availability_zone = self.availability_zone
except CloudAvailabilityZone.DoesNotExist:
availability_zone = None
if availability_zone is not None:
if availability_zone.account_id != self.account_id:
raise ValidationError("实例账号必须与可用区账号一致。")
if availability_zone.region_id != self.region_id:
raise ValidationError("实例地域必须与可用区所属地域一致。")
def __str__(self):
return "%s - %s" % (self.provider_resource_id, self.name or "未命名")
class ComputeInstanceTag(models.Model):
class Source(models.TextChoices):
PROVIDER = "provider", "云厂商"
MANUAL = "manual", "人工维护"
instance = models.ForeignKey(
ComputeInstance,
on_delete=models.CASCADE,
related_name="tags",
verbose_name="计算实例",
)
key = models.CharField("键", max_length=100)
value = models.CharField("值", max_length=255, blank=True)
source = models.CharField("来源", max_length=20, choices=Source.choices)
created_at = models.DateTimeField("创建时间", auto_now_add=True)
updated_at = models.DateTimeField("更新时间", auto_now=True)
class Meta:
ordering = ["source", "key"]
constraints = [
models.UniqueConstraint(
fields=["instance", "source", "key"],
name="cmdb_unique_instance_tag_source_key",
)
]
verbose_name = "计算实例标签"
verbose_name_plural = "计算实例标签"
def __str__(self):
return "%s=%s" % (self.key, self.value)
class SyncRun(models.Model):
class Status(models.TextChoices):
PENDING = "pending", "等待中"
RUNNING = "running", "运行中"
SUCCEEDED = "succeeded", "成功"
PARTIAL = "partial", "部分成功"
FAILED = "failed", "失败"
class Trigger(models.TextChoices):
MANUAL = "manual", "页面手动触发"
COMMAND = "command", "管理命令触发"
API = "api", "API 触发"
public_id = models.UUIDField("公开 ID", default=uuid.uuid4, unique=True, editable=False)
account = models.ForeignKey(
CloudAccount,
on_delete=models.PROTECT,
related_name="sync_runs",
verbose_name="云账号",
)
status = models.CharField(
"状态",
max_length=20,
choices=Status.choices,
default=Status.PENDING,
)
trigger = models.CharField("触发方式", max_length=20, choices=Trigger.choices)
requested_by = models.ForeignKey(
settings.AUTH_USER_MODEL,
on_delete=models.SET_NULL,
related_name="cmdb_sync_runs",
verbose_name="请求用户",
null=True,
blank=True,
)
started_at = models.DateTimeField("开始时间", null=True, blank=True)
finished_at = models.DateTimeField("结束时间", null=True, blank=True)
discovered_count = models.PositiveIntegerField("发现数量", default=0)
created_count = models.PositiveIntegerField("新建数量", default=0)
updated_count = models.PositiveIntegerField("更新数量", default=0)
unchanged_count = models.PositiveIntegerField("未变化数量", default=0)
missing_count = models.PositiveIntegerField("标记缺失数量", default=0)
restored_count = models.PositiveIntegerField("恢复数量", default=0)
retired_count = models.PositiveIntegerField("退役数量", default=0)
error_code = models.CharField("错误代码", max_length=100, blank=True)
error_message = models.TextField("错误信息", blank=True)
created_at = models.DateTimeField("创建时间", auto_now_add=True)
class Meta:
ordering = ["-created_at"]
verbose_name = "同步执行"
verbose_name_plural = "同步执行"
def __str__(self):
return "%s - %s" % (self.account.account_key, self.public_id)
class ComputeInstanceChange(models.Model):
class Action(models.TextChoices):
CREATED = "created", "创建"
UPDATED = "updated", "更新"
MARKED_MISSING = "marked_missing", "标记缺失"
RESTORED = "restored", "恢复"
RETIRED = "retired", "退役"
sync_run = models.ForeignKey(
SyncRun,
on_delete=models.PROTECT,
related_name="changes",
verbose_name="同步执行",
null=True,
blank=True,
)
instance = models.ForeignKey(
ComputeInstance,
on_delete=models.PROTECT,
related_name="changes",
verbose_name="计算实例",
)
action = models.CharField("动作", max_length=30, choices=Action.choices)
changed_fields = models.JSONField("变化字段", default=list, blank=True)
before_data = models.JSONField("变化前", default=dict, blank=True)
after_data = models.JSONField("变化后", default=dict, blank=True)
created_at = models.DateTimeField("创建时间", auto_now_add=True)
class Meta:
ordering = ["-created_at"]
verbose_name = "计算实例变化"
verbose_name_plural = "计算实例变化"
def __str__(self):
return "%s - %s" % (self.instance.provider_resource_id, self.action)
10.3.1 导入与两个账号验证器
re 验证凭据环境变量前缀;uuid 生成 SyncRun 的公开 ID。validate_credential_profile 只允许大写字母开头,随后可用大写字母、数字和下划线。它保存的是环境变量前缀,不是 AccessKey 或其他密钥。
validate_region_allowlist 要求值必须是列表;每项必须是去掉首尾空格后仍非空、长度不超过 100 的字符串,并禁止重复。模型验证用于挡住管理后台、脚本和其他绕过表单的入口;CloudAccountForm 的 clean_region_allowlist_text 在把逗号文本转为列表后还会复用同一个验证器,避免表单与模型规则漂移。
10.3.2 资产身份与关系约束
账号身份由 (provider, account_key) 唯一确定;地域、可用区和实例都以“账号 + 厂商资源 ID”唯一确定。可用区与实例的 clean() 检查账号、地域和可用区是否属于同一关系链,避免跨账号拼接。
10.3.3 运行状态与生命周期必须分开
normalized_status 表达运行中、已停止等计算状态;lifecycle_state 表达 CMDB 是否仍能在完整发现中看到它。一个实例可以是 stopped + present,也可以保留最后一次 running 状态但当前为 missing。
10.3.4 标签按来源建立独立命名空间
唯一约束是 (instance, source, key),不是 (instance, key)。因此云厂商标签 owner=provider-team 与人工标签 owner=local-team 可以同时存在,同步只接管 source=provider 的集合。
10.3.5 审计模型不保存秘密
变化快照只包含实例字段、生命周期时间和数据库关联 ID,不包含账号凭据。ProviderError 进入 SyncRun 时也只保存稳定错误码与面向操作者的脱敏消息。
10.4 数据库即时检查点
最终仓库已经包含迁移文件。先确认模型没有遗漏到迁移之外;该命令只检查,不生成文件。
执行目录:devopsX/;说明:检查最终模型与现有迁移是否一致。
python manage.py makemigrations --check --dry-run
预期结果是 No changes detected。如果出现待生成迁移,说明读者侧文件没有与本篇最终代码保持一致。
执行目录:devopsX/;说明:应用尚未执行的迁移。
python manage.py migrate
全新数据库会逐项显示以 Applying 开头并以 OK 结束的迁移结果;已经完成上一篇迁移的数据库会显示没有待应用迁移。这两种结果都正常。
执行目录:devopsX/;说明:执行 Django 系统检查。
python manage.py check
预期结果是系统检查没有发现问题。
执行目录:devopsX/;说明:运行 CMDB 模型的 7 个测试。
python manage.py test cmdb.tests.test_models.CmdbModelTests
预期结果是发现 7 个测试并以 OK 结束。
11 Provider DTO、DiscoveryScope、DiscoverySnapshot 与 ProviderError
11.1 为什么 Provider 不能直接返回 ORM 对象
云厂商 SDK 的字段命名、分页、认证与异常形式都不同。如果 Provider 直接创建 Django Model,网络层就会同时掌握数据库事务、生命周期和审计规则,任何新厂商都会复制一套核心业务逻辑。
本项目把边界固定为:Provider 只负责“读云并标准化”,服务层负责“验证并写库”。DTO 使用 @dataclass(frozen=True),调用方不能在传递过程中随意改写;元组表达快照集合,避免把可变列表当作跨层合同。
11.2 最终 Provider 基础合同
相对路径:devopsX/cmdb/providers/base.py;内容:完整文件。
from dataclasses import dataclass, field
from datetime import datetime
from typing import Protocol
@dataclass(frozen=True)
class DiscoveredRegion:
provider_resource_id: str
name: str
endpoint: str = ""
@dataclass(frozen=True)
class DiscoveredAvailabilityZone:
provider_resource_id: str
region_provider_resource_id: str
name: str
@dataclass(frozen=True)
class DiscoveredTag:
key: str
value: str
@dataclass(frozen=True)
class DiscoveredComputeInstance:
provider_resource_id: str
region_provider_resource_id: str
availability_zone_provider_resource_id: str | None = None
name: str = ""
instance_type: str = ""
vcpu: int = 0
memory_mb: int = 0
os_name: str = ""
provider_status: str = ""
normalized_status: str = "unknown"
private_ips: tuple[str, ...] = ()
public_ips: tuple[str, ...] = ()
cloud_created_at: datetime | None = None
tags: tuple[DiscoveredTag, ...] = ()
@dataclass(frozen=True)
class DiscoveryScope:
mode: str = "all"
region_ids: tuple[str, ...] = ()
@dataclass(frozen=True)
class DiscoverySnapshot:
regions: tuple[DiscoveredRegion, ...] = ()
availability_zones: tuple[DiscoveredAvailabilityZone, ...] = ()
instances: tuple[DiscoveredComputeInstance, ...] = ()
complete: bool = True
error_code: str = ""
error_message: str = ""
scope: DiscoveryScope = field(default_factory=DiscoveryScope)
class ProviderError(Exception):
def __init__(
self,
code: str,
message: str,
snapshot: DiscoverySnapshot | None = None,
):
self.code = code
self.message = message
self.snapshot = snapshot
super().__init__(message)
class CloudProviderAdapter(Protocol):
def discover(self, account) -> DiscoverySnapshot:
"""Return a normalized snapshot without returning ORM objects."""
11.2.1 四类资源 DTO
DiscoveredRegion保存厂商地域 ID、名称和可选端点。DiscoveredAvailabilityZone同时携带自己的 ID 与所属地域的厂商 ID,使服务层能建立关系。DiscoveredTag是最小键值 DTO。DiscoveredComputeInstance保存标准化实例字段;IP 与标签使用元组,云上创建时间允许为空。
11.2.2 Scope 是快照安全性的组成部分
DiscoveryScope(mode="all") 表示本次快照覆盖账号全部地域。mode="allowlist" 表示只覆盖声明的 region_ids。服务层绝不能仅看实例列表来猜范围。
| 账号配置 | 允许的快照 scope | 完整快照额外条件 |
|---|---|---|
| 配置了地域白名单 | 必须是 allowlist,且快照 region_ids 与账号白名单集合完全匹配。 | 快照 regions 必须覆盖每一个已声明地域。 |
| 没有配置地域白名单 | 只能是 all。 | 不使用 allowlist 覆盖检查。 |
这里的“完全匹配”按集合比较,顺序不影响结果;但 scope 自己仍不允许重复项或空地域 ID。对于任意 allowlist 快照,实际返回的 regions 还必须是声明 region_ids 的子集;partial 可以少返回,却不能返回范围外地域。这些合同共同阻止范围漂移导致资产误写或误判。
11.2.3 complete 决定能否根据缺席做判断
complete=True 才允许服务层把范围内“没有出现在快照中的已有资源”视为缺失。complete=False 表示部分结果:已返回数据仍可创建或更新,但不能执行缺失推断;SyncRun 最终状态为 partial。
11.2.4 ProviderError 与部分快照不是同一条路径
ProviderError 保存稳定 code、脱敏 message,并允许携带 snapshot。不过当前 sync_account 在捕获 ProviderError 时只把执行标为失败,并不会应用 exc.snapshot。因此,当前实现若希望合法写入部分结果,应直接返回 complete=False 的 DiscoverySnapshot;如果抛异常,则本次不写资产。
11.2.5 Protocol 只规定最小方法
CloudProviderAdapter 是结构化类型合同。对象只要提供兼容的 discover(account),就能作为 adapter 注入服务和测试,不要求继承某个具体基类。
12 Fake Provider:四种可重复的发现情景
12.1 Fake Provider 不是随手拼出的测试数据
Fake Provider 是可执行的 Provider 合同:它返回与真实 Provider 同一种 DTO 和 DiscoverySnapshot,也走同一套服务层验证、事务、生命周期与页面。这样可以在没有云账号、没有外网、没有真实费用的情况下练完整闭环。
12.2 最终 Fake Provider 文件
相对路径:devopsX/cmdb/providers/fake.py;内容:完整文件。
from dataclasses import replace
from datetime import datetime, timezone
from .base import (
CloudProviderAdapter,
DiscoveredAvailabilityZone,
DiscoveredComputeInstance,
DiscoveredRegion,
DiscoveredTag,
DiscoveryScope,
DiscoverySnapshot,
ProviderError,
)
class FakeProvider(CloudProviderAdapter):
scenarios = {}
def __init__(self, scenario="default"):
self.scenario = scenario
def discover(self, account):
if self.scenario not in self.scenarios:
raise ProviderError(
"UNKNOWN_FAKE_SCENARIO",
"未知 Fake Provider 场景:%s" % self.scenario,
)
scenario = self.scenarios[self.scenario]
if isinstance(scenario, Exception):
raise scenario
if callable(scenario):
scenario = scenario(account)
allowlist = tuple(account.region_allowlist)
if not allowlist:
return scenario
default_snapshot = self.scenarios["default"]
known_region_ids = {
region.provider_resource_id for region in default_snapshot.regions
}
unknown_region_ids = set(allowlist) - known_region_ids
if unknown_region_ids:
raise ProviderError(
"UNKNOWN_REGION_ALLOWLIST",
"Fake Provider 地域白名单包含未知地域:%s"
% ", ".join(sorted(unknown_region_ids)),
)
source_regions = scenario.regions
source_zones = scenario.availability_zones
if self.scenario == "empty":
source_regions = default_snapshot.regions
source_zones = default_snapshot.availability_zones
return replace(
scenario,
regions=tuple(
region
for region in source_regions
if region.provider_resource_id in allowlist
),
availability_zones=tuple(
zone
for zone in source_zones
if zone.region_provider_resource_id in allowlist
),
instances=tuple(
instance
for instance in scenario.instances
if instance.region_provider_resource_id in allowlist
),
scope=DiscoveryScope(mode="allowlist", region_ids=allowlist),
)
FakeProvider.scenarios = {
"default": DiscoverySnapshot(
regions=(
DiscoveredRegion("cn-hangzhou", "杭州", "ecs.cn-hangzhou.aliyuncs.com"),
DiscoveredRegion("cn-shanghai", "上海", "ecs.cn-shanghai.aliyuncs.com"),
),
availability_zones=(
DiscoveredAvailabilityZone("cn-hangzhou-i", "cn-hangzhou", "杭州可用区 I"),
DiscoveredAvailabilityZone("cn-shanghai-a", "cn-shanghai", "上海可用区 A"),
),
instances=(
DiscoveredComputeInstance(
provider_resource_id="i-fake-hz-001",
region_provider_resource_id="cn-hangzhou",
availability_zone_provider_resource_id="cn-hangzhou-i",
name="订单服务-演示",
instance_type="ecs.g7.large",
vcpu=2,
memory_mb=8192,
os_name="Alibaba Cloud Linux 3",
provider_status="Running",
normalized_status="running",
private_ips=("10.10.1.10",),
public_ips=("203.0.113.10",),
cloud_created_at=datetime(2026, 1, 10, tzinfo=timezone.utc),
tags=(
DiscoveredTag("department", "platform"),
DiscoveredTag("environment", "demo"),
),
),
DiscoveredComputeInstance(
provider_resource_id="i-fake-sh-001",
region_provider_resource_id="cn-shanghai",
availability_zone_provider_resource_id="cn-shanghai-a",
name="构建服务-演示",
instance_type="ecs.c7.large",
vcpu=2,
memory_mb=8192,
os_name="Ubuntu 22.04",
provider_status="Stopped",
normalized_status="stopped",
private_ips=("10.20.1.10",),
public_ips=(),
cloud_created_at=datetime(2026, 1, 11, tzinfo=timezone.utc),
tags=(DiscoveredTag("department", "devops"),),
),
),
),
"empty": DiscoverySnapshot(),
"partial": DiscoverySnapshot(
regions=(DiscoveredRegion("cn-hangzhou", "杭州"),),
instances=(),
complete=False,
error_code="PAGINATION_INCOMPLETE",
error_message="演示第二页发现失败,不能据此标记已有资产缺失。",
),
"failed": ProviderError("PROVIDER_UNAVAILABLE", "演示 Provider 暂时不可用。"),
}
12.2.1 discover 的前半段逐组解释
- 先检查场景名;未注册场景抛出
UNKNOWN_FAKE_SCENARIO。 - 场景值如果本身是异常就抛出;如果是 callable,就把 account 传入后取得快照。这给测试扩展留下了注入点。
- 账号没有白名单时直接返回场景快照,默认 scope 仍为
all。 - 账号有白名单时,用 default 场景中的地域集合判断哪些地域是 Fake Provider 已知地域。
- 出现未知地域时抛出
UNKNOWN_REGION_ALLOWLIST,不会悄悄返回空快照。
12.2.2 allowlist 过滤的关键细节
普通场景按地域过滤 regions、zones 和 instances,并把 scope 改成与账号配置一致的 allowlist。empty 是一个特殊但必要的分支:它虽然没有实例,却从 default 场景取地域和可用区再按白名单过滤。这样“完整空实例快照”仍覆盖每一个声明地域,既满足最新 scope 合同,又能只对范围内实例做缺失判断。
12.2.3 四个内置场景
| 场景 | 返回内容 | complete | 服务层结果 |
|---|---|---|---|
| default | 杭州、上海各一个可用区与一台实例。 | true | 首次创建两台;重复运行计为 unchanged。 |
| empty | 无实例;无白名单时资源集合全空,有白名单时仍返回范围地域与可用区。 | true | 把范围内未返回实例标记为 missing。 |
| partial | 只包含杭州地域,无实例,并带分页未完成错误。 | false | 状态为 partial,绝不根据缺席标记 missing。 |
| failed | 直接抛 PROVIDER_UNAVAILABLE。 | 无快照 | SyncRun 失败,资产不变。 |
12.3 Fake allowlist 合同的即时测试
执行目录:devopsX/;说明:验证已配置白名单时只发现范围内资源。
python manage.py test cmdb.tests.test_sync.SyncAccountTests.test_fake_provider_filters_configured_region_allowlist
预期发现 1 个测试并以 OK 结束;断言最终只存在杭州实例和杭州地域。
执行目录:devopsX/;说明:验证 Fake Provider 拒绝未知白名单地域。
python manage.py test cmdb.tests.test_sync.SyncAccountTests.test_fake_provider_rejects_unknown_allowlist_region
预期以 OK 结束;测试内部确认错误码为 UNKNOWN_REGION_ALLOWLIST,对应 SyncRun 为 failed。
执行目录:devopsX/;说明:验证未注册 Fake 场景被明确拒绝。
python manage.py test cmdb.tests.test_sync.SyncAccountTests.test_unknown_fake_scenario_is_rejected
预期以 OK 结束;管理命令本身还通过 choices 限制场景,直接注入 adapter 时则由 Provider 再防守一次。
13 sync_account 服务流:短锁、事务边界与网络调用
13.1 服务层是唯一对账入口
页面、管理命令和将来的 API 都调用 sync_account。入口统一后,禁用检查、并发保护、验证、审计、生命周期和错误脱敏才不会在不同入口中漂移。
13.2 最终同步服务文件
相对路径:devopsX/cmdb/services/sync.py;内容:完整文件。
import logging
from datetime import timedelta
from django.db import connection, transaction
from django.db.models import F, Q
from django.utils import timezone
from cmdb.models import (
CloudAccount,
CloudAvailabilityZone,
CloudRegion,
ComputeInstance,
ComputeInstanceChange,
ComputeInstanceTag,
SyncRun,
)
from ..providers.base import DiscoverySnapshot, ProviderError
from ..providers.fake import FakeProvider
logger = logging.getLogger(__name__)
MANAGED_FIELDS = (
"region_id",
"availability_zone_id",
"name",
"instance_type",
"vcpu",
"memory_mb",
"os_name",
"provider_status",
"normalized_status",
"private_ips",
"public_ips",
"cloud_created_at",
"provider_tags",
)
SYNC_RUN_STALE_AFTER = timedelta(minutes=30)
def _lock_account(account_id):
if connection.vendor == "sqlite":
CloudAccount.objects.filter(pk=account_id).update(
updated_at=F("updated_at")
)
return (
CloudAccount.objects.select_for_update()
.select_related("provider")
.get(pk=account_id)
)
def _sync_configuration(account):
return (
account.provider_id,
account.provider.code,
account.provider.is_active,
account.is_active,
account.sync_enabled,
account.credential_profile,
tuple(sorted(account.region_allowlist)),
)
def get_provider_adapter(account, scenario="default"):
if account.provider.code == "fake":
return FakeProvider(scenario=scenario)
if account.provider.code == "aliyun":
from ..providers.aliyun import AliyunEcsProvider
return AliyunEcsProvider()
raise ProviderError("UNSUPPORTED_PROVIDER", "暂不支持云厂商:%s" % account.provider.code)
def _instance_state(instance):
return {
"region_id": instance.region_id,
"availability_zone_id": instance.availability_zone_id,
"name": instance.name,
"instance_type": instance.instance_type,
"vcpu": instance.vcpu,
"memory_mb": instance.memory_mb,
"os_name": instance.os_name,
"provider_status": instance.provider_status,
"normalized_status": instance.normalized_status,
"private_ips": instance.private_ips,
"public_ips": instance.public_ips,
"cloud_created_at": instance.cloud_created_at.isoformat()
if instance.cloud_created_at
else None,
"lifecycle_state": instance.lifecycle_state,
"missing_since": (
instance.missing_since.isoformat() if instance.missing_since else None
),
"retired_at": instance.retired_at.isoformat() if instance.retired_at else None,
"provider_tags": {
key: value
for key, value in instance.tags.filter(
source=ComputeInstanceTag.Source.PROVIDER
).order_by("key").values_list("key", "value")
},
}
def _validate_snapshot(account, snapshot):
if snapshot.scope.mode not in {"all", "allowlist"}:
raise ProviderError(
"INVALID_DISCOVERY_SCOPE",
"发现范围模式无效:%s" % snapshot.scope.mode,
)
if snapshot.scope.mode == "allowlist":
if not snapshot.scope.region_ids:
raise ProviderError(
"INVALID_DISCOVERY_SCOPE",
"地域白名单发现范围不能为空。",
)
if len(set(snapshot.scope.region_ids)) != len(snapshot.scope.region_ids):
raise ProviderError(
"INVALID_DISCOVERY_SCOPE",
"地域白名单发现范围包含重复地域 ID。",
)
if any(not region_id for region_id in snapshot.scope.region_ids):
raise ProviderError(
"INVALID_DISCOVERY_SCOPE",
"地域白名单发现范围包含空地域 ID。",
)
configured_allowlist = tuple(account.region_allowlist)
if configured_allowlist:
if snapshot.scope.mode != "allowlist":
raise ProviderError(
"INVALID_DISCOVERY_SCOPE",
"云账号配置了地域白名单,但快照未声明白名单范围。",
)
if set(snapshot.scope.region_ids) != set(configured_allowlist):
raise ProviderError(
"INVALID_DISCOVERY_SCOPE",
"发现快照范围与云账号地域白名单不一致。",
)
elif snapshot.scope.mode != "all":
raise ProviderError(
"INVALID_DISCOVERY_SCOPE",
"未配置地域白名单的云账号只能接受全部地域范围的快照。",
)
region_ids = set()
for region in snapshot.regions:
if not region.provider_resource_id:
raise ProviderError("EMPTY_REGION_ID", "发现了空地域 ID。")
if region.provider_resource_id in region_ids:
raise ProviderError("DUPLICATE_REGION", "发现重复地域 ID:%s" % region.provider_resource_id)
region_ids.add(region.provider_resource_id)
if snapshot.scope.mode == "allowlist":
declared_region_ids = set(snapshot.scope.region_ids)
if not region_ids.issubset(declared_region_ids):
raise ProviderError(
"OUT_OF_SCOPE_DISCOVERY_REGION",
"发现快照包含地域白名单范围外的地域。",
)
if snapshot.complete and declared_region_ids != region_ids:
raise ProviderError(
"INCOMPLETE_DISCOVERY_SCOPE",
"完整快照没有覆盖地域白名单中的全部地域。",
)
zone_ids = set()
zone_regions = {}
for zone in snapshot.availability_zones:
if not zone.provider_resource_id:
raise ProviderError("EMPTY_ZONE_ID", "发现了空可用区 ID。")
if zone.provider_resource_id in zone_ids:
raise ProviderError("DUPLICATE_ZONE", "发现重复可用区 ID:%s" % zone.provider_resource_id)
if zone.region_provider_resource_id not in region_ids:
raise ProviderError(
"UNKNOWN_ZONE_REGION",
"可用区 %s 引用了未知地域 %s"
% (zone.provider_resource_id, zone.region_provider_resource_id),
)
zone_ids.add(zone.provider_resource_id)
zone_regions[zone.provider_resource_id] = zone.region_provider_resource_id
valid_statuses = set(ComputeInstance.NormalizedStatus.values)
instance_ids = set()
for instance in snapshot.instances:
if not instance.provider_resource_id:
raise ProviderError("EMPTY_INSTANCE_ID", "发现了空实例 ID。")
if instance.provider_resource_id in instance_ids:
raise ProviderError(
"DUPLICATE_INSTANCE",
"发现重复实例 ID:%s" % instance.provider_resource_id,
)
if instance.region_provider_resource_id not in region_ids:
raise ProviderError(
"UNKNOWN_INSTANCE_REGION",
"实例 %s 引用了未知地域 %s"
% (instance.provider_resource_id, instance.region_provider_resource_id),
)
zone_id = instance.availability_zone_provider_resource_id
if zone_id and zone_id not in zone_ids:
raise ProviderError(
"UNKNOWN_INSTANCE_ZONE",
"实例 %s 引用了未知可用区 %s"
% (instance.provider_resource_id, zone_id),
)
if zone_id and zone_regions[zone_id] != instance.region_provider_resource_id:
raise ProviderError(
"INSTANCE_ZONE_REGION_MISMATCH",
"实例 %s 的地域与可用区所属地域不一致。"
% instance.provider_resource_id,
)
if instance.normalized_status not in valid_statuses:
raise ProviderError(
"INVALID_NORMALIZED_STATUS",
"实例 %s 的标准状态无效:%s"
% (instance.provider_resource_id, instance.normalized_status),
)
instance_ids.add(instance.provider_resource_id)
def _upsert_region(account, item, now):
region, created = CloudRegion.objects.get_or_create(
account=account,
provider_resource_id=item.provider_resource_id,
defaults={
"name": item.name,
"endpoint": item.endpoint,
"first_seen_at": now,
"last_seen_at": now,
"is_active": True,
},
)
if not created:
region.name = item.name
region.endpoint = item.endpoint
region.is_active = True
region.last_seen_at = now
region.save(update_fields=["name", "endpoint", "is_active", "last_seen_at"])
return region
def _upsert_zone(account, region_map, item, now):
region = region_map[item.region_provider_resource_id]
zone, created = CloudAvailabilityZone.objects.get_or_create(
account=account,
provider_resource_id=item.provider_resource_id,
defaults={
"region": region,
"name": item.name,
"first_seen_at": now,
"last_seen_at": now,
"is_active": True,
},
)
if not created:
zone.region = region
zone.name = item.name
zone.is_active = True
zone.last_seen_at = now
zone.save(update_fields=["region", "name", "is_active", "last_seen_at"])
return zone
def _provider_tag_state(tags):
return {tag.key: tag.value for tag in sorted(tags, key=lambda tag: tag.key)}
def _upsert_provider_tags(instance, tags):
incoming = {tag.key: tag.value for tag in tags}
existing = set(
ComputeInstanceTag.objects.filter(
instance=instance,
source=ComputeInstanceTag.Source.PROVIDER,
).values_list("key", flat=True)
)
for key in existing - set(incoming):
ComputeInstanceTag.objects.filter(
instance=instance,
source=ComputeInstanceTag.Source.PROVIDER,
key=key,
).delete()
for key, value in incoming.items():
ComputeInstanceTag.objects.update_or_create(
instance=instance,
source=ComputeInstanceTag.Source.PROVIDER,
key=key,
defaults={"value": value},
)
def _apply_snapshot(account, sync_run, snapshot):
now = timezone.now()
region_map = {}
for item in snapshot.regions:
region_map[item.provider_resource_id] = _upsert_region(account, item, now)
zone_map = {}
for item in snapshot.availability_zones:
zone_map[item.provider_resource_id] = _upsert_zone(account, region_map, item, now)
discovered_ids = set()
created_count = 0
updated_count = 0
unchanged_count = 0
restored_count = 0
for item in snapshot.instances:
region = region_map[item.region_provider_resource_id]
zone = zone_map.get(item.availability_zone_provider_resource_id)
defaults = {
"region": region,
"availability_zone": zone,
"name": item.name,
"instance_type": item.instance_type,
"vcpu": item.vcpu,
"memory_mb": item.memory_mb,
"os_name": item.os_name,
"provider_status": item.provider_status,
"normalized_status": item.normalized_status,
"lifecycle_state": ComputeInstance.LifecycleState.PRESENT,
"private_ips": list(item.private_ips),
"public_ips": list(item.public_ips),
"cloud_created_at": item.cloud_created_at,
"first_seen_at": now,
"last_seen_at": now,
"missing_since": None,
"retired_at": None,
}
instance = ComputeInstance.objects.filter(
account=account,
provider_resource_id=item.provider_resource_id,
).first()
if instance is None:
instance = ComputeInstance.objects.create(
account=account,
provider_resource_id=item.provider_resource_id,
**defaults,
)
_upsert_provider_tags(instance, item.tags)
ComputeInstanceChange.objects.create(
sync_run=sync_run,
instance=instance,
action=ComputeInstanceChange.Action.CREATED,
changed_fields=list(MANAGED_FIELDS),
before_data={},
after_data=_instance_state(instance),
)
created_count += 1
else:
before = _instance_state(instance)
was_missing = instance.lifecycle_state == ComputeInstance.LifecycleState.MISSING
was_retired = instance.lifecycle_state == ComputeInstance.LifecycleState.RETIRED
after = dict(before)
after.update(
{
"region_id": region.id,
"availability_zone_id": zone.id if zone else None,
"name": item.name,
"instance_type": item.instance_type,
"vcpu": item.vcpu,
"memory_mb": item.memory_mb,
"os_name": item.os_name,
"provider_status": item.provider_status,
"normalized_status": item.normalized_status,
"private_ips": list(item.private_ips),
"public_ips": list(item.public_ips),
"cloud_created_at": item.cloud_created_at.isoformat()
if item.cloud_created_at
else None,
"provider_tags": _provider_tag_state(item.tags),
}
)
changed_fields = [field for field in MANAGED_FIELDS if before[field] != after[field]]
instance.region = region
instance.availability_zone = zone
instance.name = item.name
instance.instance_type = item.instance_type
instance.vcpu = item.vcpu
instance.memory_mb = item.memory_mb
instance.os_name = item.os_name
instance.provider_status = item.provider_status
instance.normalized_status = item.normalized_status
if not was_retired:
instance.lifecycle_state = ComputeInstance.LifecycleState.PRESENT
instance.private_ips = list(item.private_ips)
instance.public_ips = list(item.public_ips)
instance.cloud_created_at = item.cloud_created_at
instance.last_seen_at = now
if not was_retired:
instance.missing_since = None
instance.retired_at = None
instance.save()
_upsert_provider_tags(instance, item.tags)
current_state = _instance_state(instance)
if was_missing:
ComputeInstanceChange.objects.create(
sync_run=sync_run,
instance=instance,
action=ComputeInstanceChange.Action.RESTORED,
changed_fields=changed_fields
+ ["lifecycle_state", "missing_since"],
before_data=before,
after_data=current_state,
)
restored_count += 1
elif changed_fields:
ComputeInstanceChange.objects.create(
sync_run=sync_run,
instance=instance,
action=ComputeInstanceChange.Action.UPDATED,
changed_fields=changed_fields,
before_data=before,
after_data=current_state,
)
updated_count += 1
else:
unchanged_count += 1
discovered_ids.add(item.provider_resource_id)
missing_count = 0
if snapshot.complete:
seen_region_ids = {item.provider_resource_id for item in snapshot.regions}
seen_zone_ids = {
item.provider_resource_id for item in snapshot.availability_zones
}
scope_mode = snapshot.scope.mode
scoped_region_ids = snapshot.scope.region_ids
region_scope = CloudRegion.objects.filter(account=account)
zone_scope = CloudAvailabilityZone.objects.filter(account=account)
active_instances = ComputeInstance.objects.filter(account=account).exclude(
lifecycle_state=ComputeInstance.LifecycleState.RETIRED
)
if scope_mode == "allowlist":
region_scope = region_scope.filter(
provider_resource_id__in=scoped_region_ids
)
zone_scope = zone_scope.filter(
region__provider_resource_id__in=scoped_region_ids
)
active_instances = active_instances.filter(
region__provider_resource_id__in=scoped_region_ids
)
region_scope.exclude(provider_resource_id__in=seen_region_ids).update(
is_active=False
)
zone_scope.exclude(provider_resource_id__in=seen_zone_ids).update(
is_active=False
)
for instance in active_instances:
if instance.provider_resource_id in discovered_ids:
continue
if instance.lifecycle_state == ComputeInstance.LifecycleState.MISSING:
continue
before = _instance_state(instance)
instance.lifecycle_state = ComputeInstance.LifecycleState.MISSING
instance.missing_since = instance.missing_since or now
instance.save(update_fields=["lifecycle_state", "missing_since", "updated_at"])
ComputeInstanceChange.objects.create(
sync_run=sync_run,
instance=instance,
action=ComputeInstanceChange.Action.MARKED_MISSING,
changed_fields=["lifecycle_state", "missing_since"],
before_data=before,
after_data=_instance_state(instance),
)
missing_count += 1
return {
"discovered_count": len(snapshot.instances),
"created_count": created_count,
"updated_count": updated_count,
"unchanged_count": unchanged_count,
"missing_count": missing_count,
"restored_count": restored_count,
"retired_count": 0,
}
def sync_account(
account,
requested_by=None,
trigger=SyncRun.Trigger.MANUAL,
scenario="default",
adapter=None,
):
with transaction.atomic():
locked_account = _lock_account(account.pk)
if not locked_account.provider.is_active:
raise ProviderError("PROVIDER_DISABLED", "云厂商已停用。")
if not locked_account.is_active or not locked_account.sync_enabled:
raise ProviderError("ACCOUNT_DISABLED", "云账号未启用同步。")
now = timezone.now()
stale_before = now - SYNC_RUN_STALE_AFTER
SyncRun.objects.filter(
Q(started_at__lt=stale_before)
| Q(started_at__isnull=True, created_at__lt=stale_before),
account=locked_account,
status__in=[SyncRun.Status.PENDING, SyncRun.Status.RUNNING],
).update(
status=SyncRun.Status.FAILED,
error_code="STALE_SYNC_RUN",
error_message="上一次同步超过 30 分钟未结束,已自动解除阻塞。",
finished_at=now,
)
if SyncRun.objects.filter(
account=locked_account,
status__in=[SyncRun.Status.PENDING, SyncRun.Status.RUNNING],
).exists():
raise ProviderError("SYNC_ALREADY_RUNNING", "该云账号已有同步正在执行。")
discovery_configuration = _sync_configuration(locked_account)
sync_run = SyncRun.objects.create(
account=locked_account,
status=SyncRun.Status.RUNNING,
trigger=trigger,
requested_by=requested_by,
started_at=now,
)
try:
adapter = adapter or get_provider_adapter(locked_account, scenario=scenario)
snapshot = adapter.discover(locked_account)
_validate_snapshot(locked_account, snapshot)
with transaction.atomic():
locked_account = _lock_account(locked_account.pk)
current_run = SyncRun.objects.select_for_update().get(pk=sync_run.pk)
if current_run.status != SyncRun.Status.RUNNING:
raise ProviderError(
"SYNC_RUN_SUPERSEDED",
"本次同步已被更新的同步执行取代,快照未写入。",
)
if _sync_configuration(locked_account) != discovery_configuration:
raise ProviderError(
"SYNC_CONFIGURATION_CHANGED",
"同步期间云账号或云厂商配置发生变化,快照未写入。",
)
_validate_snapshot(locked_account, snapshot)
stats = _apply_snapshot(locked_account, current_run, snapshot)
current_run.status = (
SyncRun.Status.SUCCEEDED
if snapshot.complete
else SyncRun.Status.PARTIAL
)
current_run.error_code = snapshot.error_code
current_run.error_message = snapshot.error_message
for key, value in stats.items():
setattr(current_run, key, value)
current_run.finished_at = timezone.now()
current_run.save()
if snapshot.complete:
locked_account.last_successful_sync_at = current_run.finished_at
locked_account.save(update_fields=["last_successful_sync_at", "updated_at"])
sync_run = current_run
return sync_run
except ProviderError as exc:
if exc.code.startswith("ALIYUN_"):
logger.warning(
"CMDB synchronization failed code=%s account_id=%s sync_run_id=%s.",
exc.code,
locked_account.pk,
sync_run.pk,
)
with transaction.atomic():
SyncRun.objects.filter(
pk=sync_run.pk,
status=SyncRun.Status.RUNNING,
).update(
status=SyncRun.Status.FAILED,
error_code=exc.code,
error_message=exc.message,
finished_at=timezone.now(),
)
sync_run.refresh_from_db()
raise
except Exception as exc:
logger.error(
"Unexpected CMDB synchronization failure type=%s account_id=%s sync_run_id=%s.",
type(exc).__name__,
locked_account.pk,
sync_run.pk,
)
with transaction.atomic():
SyncRun.objects.filter(
pk=sync_run.pk,
status=SyncRun.Status.RUNNING,
).update(
status=SyncRun.Status.FAILED,
error_code="UNEXPECTED_ERROR",
error_message="同步失败,请查看服务日志。",
finished_at=timezone.now(),
)
sync_run.refresh_from_db()
raise ProviderError("UNEXPECTED_ERROR", "同步失败,请查看服务日志。") from exc
def retire_instance(instance, sync_run=None):
with transaction.atomic():
locked_account = _lock_account(instance.account_id)
locked_instance = ComputeInstance.objects.select_for_update().get(
pk=instance.pk,
account=locked_account,
)
if locked_instance.lifecycle_state == ComputeInstance.LifecycleState.RETIRED:
return False
before = _instance_state(locked_instance)
changed_fields = ["lifecycle_state", "retired_at"]
if locked_instance.missing_since is not None:
changed_fields.append("missing_since")
locked_instance.lifecycle_state = ComputeInstance.LifecycleState.RETIRED
locked_instance.missing_since = None
locked_instance.retired_at = timezone.now()
locked_instance.save(
update_fields=[
"lifecycle_state",
"missing_since",
"retired_at",
"updated_at",
]
)
ComputeInstanceChange.objects.create(
sync_run=sync_run,
instance=locked_instance,
action=ComputeInstanceChange.Action.RETIRED,
changed_fields=changed_fields,
before_data=before,
after_data=_instance_state(locked_instance),
)
return True
13.2.1 常量与账号锁
MANAGED_FIELDS 决定哪些 Provider 字段参与“是否发生业务更新”的比较,其中包括排序后的 provider_tags;生命周期字段单独处理。SYNC_RUN_STALE_AFTER 固定为 30 分钟。
_lock_account 在 MySQL 路径使用 select_for_update();SQLite 路径先执行一次值不变的 UPDATE,再读取账号。后文会准确说明两者差异。
13.2.2 Provider 分派与状态序列化
get_provider_adapter 按 provider.code 选择 Fake 或 Aliyun,未知厂商抛 UNSUPPORTED_PROVIDER。Aliyun 类采用函数内导入,使没有安装可选 SDK 时仍能运行 Fake 教程。
_instance_state 把实例转换成可放入 JSONField 的字典。datetime 统一转 ISO 字符串;Provider 标签按 key 排序后转成字典。这个状态既用于前后比较,也用于 ComputeInstanceChange,因此只改 Provider 标签也会形成可审计 UPDATED。
13.2.3 快照验证发生在资产写入前
_validate_snapshot 先验证 scope,再验证地域、可用区与实例的引用完整性、唯一性和标准状态。网络返回后先验证一次;第二段事务重新读取并锁住账号、确认配置未变后再验证一次。通过这两层检查前不会调用 _apply_snapshot,因此坏快照或基于旧配置取得的快照不会写入资产。
13.2.4 地域、可用区、标签与实例的对账
_upsert_region 与 _upsert_zone 使用稳定资源 ID 查找;已存在时刷新字段、last_seen 和 is_active。_upsert_provider_tags 只查询 source=provider:删除云端已经消失的 Provider 标签,并 update_or_create 当前 Provider 标签。新实例先写 Provider 标签再生成 CREATED 的 after_data;已有实例先比较归一化标签状态,再写标签并生成 UPDATED 或 RESTORED,因此标签变化不会漏出审计。
_apply_snapshot 先建立 region_map、zone_map,再处理实例。新实例写 CREATED;已有实例比较 MANAGED_FIELDS,写 UPDATED、RESTORED 或计入 unchanged。只有 complete 快照才会进入缺失扫描。
13.2.5 sync_account 的两段数据库事务
第一段事务创建受保护的运行记录,同时保存 Provider、启用状态、凭据前缀和排序后地域白名单组成的同步配置指纹;外部网络调用在事务外。第二段事务重新读取配置,只有运行仍有效且配置与发现开始时一致,才原子应用快照。异常处理再用独立短事务把仍处于 running 的执行标为 failed。
13.2.6 retire_instance 是显式人工动作
退役不是“连续若干次 missing 后自动删除”。retire_instance 自己开启事务,按与同步相同的顺序先锁账号,再锁实例;随后设置 retired_at、保留实例与历史,并写 RETIRED 变化。重复退役返回 False,不制造重复记录。这个顺序减少人工退役与同步写同一实例时的竞态。
13.3 两段短事务的真实时间线
| 阶段 | 是否在数据库事务内 | 执行内容 | 为何这样设计 |
|---|---|---|---|
| A | 是 | 锁账号;检查 Provider 与账号状态;回收 stale run;拒绝新鲜并发 run;保存同步配置;创建 running SyncRun。 | 快速建立“该账号已有一次同步”及“发现基于哪份配置”的事实。 |
| B | 否 | 选择 adapter,执行 discover,验证 DTO 快照。 | 云 API 可能慢、分页或超时,不能长时间占着数据库事务与行锁。 |
| C | 是 | 再次锁账号和当前 SyncRun;检查运行状态与同步配置;再次验证快照;原子写地域、可用区、实例、标签、变化与统计。 | 旧执行或旧配置快照被 fencing;合法快照要么整体提交,要么整体回滚。 |
| D | 异常时另开短事务 | 仅当原 SyncRun 仍为 running 时把它标为 failed。 | 不覆盖 stale recovery 或替代执行已经写入的新状态。 |
最重要的一点是:网络调用在两段事务之间。它既没有持有 MySQL 行锁,也没有把 SQLite 写锁保持到云端分页完成。
13.4 MySQL 与 SQLite 锁语义不能混为一谈
13.4.1 MySQL 路径
在支持行级锁的 MySQL 事务中,select_for_update() 会锁住读取到的 CloudAccount 行,直到当前短事务提交或回滚。第一段事务可串行化“检查是否已有活动 run + 创建新 run”;第二段事务可串行化“确认 run + 写入快照”。这是真实数据库行锁语义。
13.4.2 SQLite 路径
Django 在 SQLite 上的 select_for_update() 不会提供等价的行锁效果。当前 _lock_account 因此先执行 updated_at=F("updated_at") 的值不变 UPDATE,用一次写操作尽早取得 SQLite 写事务所需的锁,再读取账号。
这只是教学与本地开发缓解措施:SQLite 的写并发粒度更粗,竞争者可能等待或遇到数据库忙;它不是行锁,不是分布式锁,也不能证明多节点生产并发安全。设置中的 20 秒 timeout 只是等待上限,不会把 SQLite 变成 MySQL。生产并发语义应使用真实支持 SELECT FOR UPDATE 的数据库,并另外评估任务队列、超时与部署拓扑。
13.4.3 当前实现没有声称什么
- 没有声称实现 Redis/Etcd 等分布式锁。
- 没有在网络调用期间长期持锁。
- 没有把“同账号只允许一个新鲜 SyncRun”扩大成跨系统的全局互斥保证。
- 没有用本地 SQLite 测试替代 MySQL 并发测试。
14 验证、原子 upsert、标签隔离与幂等对账
14.1 快照验证规则按顺序理解
SyncRun 在网络前已经创建,所以“验证前不写入”准确地说是“不写入资产数据”;验证失败会保留一条 failed SyncRun 供审计。当前验证顺序如下。
| 层次 | 检查 | 失败风险被阻止 |
|---|---|---|
| scope 模式 | 只允许 all 或 allowlist。 | 未知范围语义。 |
| allowlist 自身 | 非空、无重复、无空 ID。 | 模糊或自相矛盾的声明。 |
| 账号与 scope | 配置白名单必须收到集合匹配的 allowlist;未配置白名单只能收到 all。 | 范围缩小导致范围外资产被误判。 |
| 地域 | ID 非空且唯一。 | 无法建立稳定身份或 map 被覆盖。 |
| allowlist 范围边界 | 无论 complete 还是 partial,快照 regions 都只能是 scope.region_ids 的子集。 | 白名单外数据越界进入账号快照。 |
| 完整 allowlist | regions 覆盖 scope 声明的每一个地域。 | 声称完整却漏掉某个白名单地域。 |
| 可用区 | ID 非空唯一,所属地域必须存在于本快照。 | 悬空外键。 |
| 实例 | ID 非空唯一,地域已知;可用区已知且属于同一地域。 | 跨地域错误关联。 |
| 标准状态 | 必须属于 ComputeInstance.NormalizedStatus。 | 页面和查询出现未定义状态。 |
14.1.1 最新 scope 合同必须逐字落实为三条
- 账号配置了 allowlist,就必须收到
mode=allowlist且 region_ids 集合与配置匹配的快照。 - 账号没有配置 allowlist,就只接受
mode=all。 - 当 allowlist 快照声明
complete=True时,快照 regions 必须覆盖每一个已声明地域。
第三条只针对 complete 快照。partial allowlist 允许尚未发现完全部声明地域,因为它本来就明确表示结果不完整;与此同时,partial 永远不能标记缺失。无论快照是否完整,只要 scope 是 allowlist,快照实际返回的 regions 都必须是已声明 region_ids 的子集;范围外地域会触发 OUT_OF_SCOPE_DISCOVERY_REGION。
14.2 原子 upsert 的准确含义
这里的“upsert”是服务层的查找后创建/更新流程,不是 MySQL 单条重复键更新语句。地域和可用区使用 get_or_create,实例使用 filter-first 后 create 或 update;数据库唯一约束仍是最后防线。
原子性来自第二段 transaction.atomic():地域、可用区、实例、Provider 标签、变化记录、SyncRun 统计和账号最后成功同步时间作为一个事务提交。若中途抛异常,这批资产写入整体回滚,随后异常处理在另一个短事务中把运行标为 failed。
这并不等于无限吞吐的批量同步。当前实现按资源逐项 ORM 操作,优先展示正确边界与审计行为;规模扩大后才能在保持合同的前提下评估 bulk 操作。
14.3 Provider 标签与人工标签互不越权
- 同步先把 incoming Provider 标签转成
{key: value}。 - 只读取 source=provider 的现有键。
- 只删除“Provider 原有、当前快照已没有”的 Provider 标签。
- 只 update_or_create source=provider 的当前标签。
- source=manual 的标签从未进入上述查询,因此不会被云同步删除或覆盖。
- Provider 标签先归一化为按 key 比较的字典;即使其他实例字段都没变,标签集合或值变化也会写 UPDATED,changed_fields 包含 provider_tags。
同一个 key 可以在两个 source 下各保存一条,这就是模型唯一约束包含 source 的原因。
14.4 幂等性要看业务结果,不要误解成零写入
第一次 default 快照创建两台实例和两条 CREATED 变化,after_data 已包含 Provider 标签。第二次完全相同的快照不会重复创建实例,不会写 UPDATED 变化,unchanged_count=2。若只改变 Provider 标签,则会产生一条 UPDATED,而不是误计 unchanged。不过完全相同的运行仍会刷新 last_seen_at、保存实例并执行标签 update_or_create;因此本项目保证的是资产身份和变化审计的业务幂等,不承诺没有 SQL。
完整空快照重复执行也具有幂等性:第一次把 present 变成 missing 并写 MARKED_MISSING;第二次看到实例已经 missing 后直接跳过,不重复写缺失变化,也不重置 missing_since。
14.5 complete 与 partial 的落库差异
| 行为 | complete | partial |
|---|---|---|
| 创建快照中出现的新实例 | 会 | 会 |
| 更新快照中出现的已有实例 | 会 | 会 |
| 恢复快照中重新出现的 missing 实例 | 会 | 会 |
| 把未出现实例标记 missing | 会,且只在 scope 内 | 绝不会 |
| 停用未出现地域/可用区 | 会,且只在 scope 内 | 不会 |
| SyncRun 状态 | succeeded | partial |
| 更新 last_successful_sync_at | 会 | 不会 |
14.6 验证与幂等即时检查点
执行目录:devopsX/;说明:验证重复地域在资产写入前被拒绝。
python manage.py test cmdb.tests.test_sync.SyncAccountTests.test_duplicate_region_is_rejected_before_writing_snapshot
预期以 OK 结束;测试确认地域数量仍为 0,并保留 failed SyncRun。
执行目录:devopsX/;说明:验证空资源身份与无效标准状态。
python manage.py test cmdb.tests.test_sync.SyncAccountTests.test_empty_resource_identity_and_invalid_status_are_rejected
预期以 OK 结束。
执行目录:devopsX/;说明:验证相同快照不重复生成变化。
python manage.py test cmdb.tests.test_sync.SyncAccountTests.test_identical_snapshot_is_idempotent
预期以 OK 结束;第二次执行 created=0、updated=0、unchanged=2。
执行目录:devopsX/;说明:验证 Provider 同步不删除人工标签。
python manage.py test cmdb.tests.test_sync.SyncAccountTests.test_provider_sync_does_not_delete_manual_tags
预期以 OK 结束;测试同时确认已从 Provider 消失的 Provider 标签会被删除。
执行目录:devopsX/;说明:验证仅 Provider 标签变化也生成 UPDATED 审计。
python manage.py test cmdb.tests.test_sync.SyncAccountTests.test_provider_tag_only_change_is_audited_as_update
预期以 OK 结束;changed_fields 包含 provider_tags。
执行目录:devopsX/;说明:验证部分快照绝不推断缺失。
python manage.py test cmdb.tests.test_sync.SyncAccountTests.test_partial_snapshot_never_marks_existing_instances_missing
预期以 OK 结束;SyncRun 为 partial,missing_count=0。
执行目录:devopsX/;说明:验证 partial allowlist 也不能包含范围外地域。
python manage.py test cmdb.tests.test_sync.SyncAccountTests.test_partial_allowlist_snapshot_rejects_out_of_scope_region
预期以 OK 结束;错误码为 OUT_OF_SCOPE_DISCOVERY_REGION,且不写地域。
15 missing、restored、retired、stale recovery 与 fencing
15.1 实例生命周期状态机
| 原状态 | 事件 | 新状态 | 变化动作 | 时间字段 |
|---|---|---|---|---|
| 不存在 | 快照首次发现 | present | CREATED | first_seen_at 与 last_seen_at 同时写入。 |
| present | 完整快照在有效 scope 内未发现 | missing | MARKED_MISSING | missing_since 首次写入。 |
| missing | 下一次仍未发现 | missing | 无新变化 | 保留原 missing_since。 |
| missing | 任一合法快照再次发现 | present | RESTORED | missing_since 清空,last_seen_at 刷新。 |
| present 或 missing | 人工退役 | retired | RETIRED | retired_at 写入;若原状态为 missing,missing_since 清空,并在 RETIRED 的 before_data、after_data 和 changed_fields 中审计。 |
| retired | 以后快照再次发现 | retired | 不会写 RESTORED;受管字段变化时仍可写 UPDATED | retired_at 保留;Provider 字段与 last_seen 仍可刷新;同步不会静默改写 retired 行已有的 missing_since。 |
退役是人工决定,当前实现不会因为资源重新出现就自动撤销,也不会删除历史。受控退役会清理旧 missing_since 并把该变化写入审计;同步重新发现 retired 行时不会进入自动恢复分支,也不会静默清空该行已有的 missing_since。若业务需要“取消退役”,应设计显式权限、动作与审计,不能悄悄复用 restored。
15.2 地域和可用区使用 is_active,而不是实例生命周期
完整快照会把 scope 内未出现的地域与可用区更新为 is_active=False;重新发现时 upsert 会恢复为 True。partial 不做停用。实例需要更细的 missing、restored 与 retired 审计,所以使用独立 lifecycle_state。
15.3 stale-run recovery 解除遗留阻塞
第一段事务会查找同账号中 pending 或 running 的执行。若 started_at 早于当前时间 30 分钟,或 pending 没有 started_at 且 created_at 已超过 30 分钟,就批量更新为 failed:
error_code=STALE_SYNC_RUN;- 错误消息说明上一次同步超过 30 分钟;
finished_at记录回收时间。
回收后再检查是否仍有新鲜 pending/running。若有,则抛 SYNC_ALREADY_RUNNING。因此 stale recovery 处理的是“旧执行遗留状态”,不是任意抢占正在正常运行的任务。
15.4 fencing 阻止旧执行与旧配置快照落库
15.4.1 运行状态 fencing
设想执行 A 网络调用很慢,超过 30 分钟;执行 B 启动时把 A 标为 stale,并成功写入新快照。随后 A 才返回。如果 A 直接落库,就可能用旧数据覆盖 B。
当前实现第二段事务重新锁定 A 的 SyncRun,并要求状态仍为 running。A 已被 B 的 stale recovery 改成 failed,因此触发 SYNC_RUN_SUPERSEDED,旧快照完全不写。
15.4.2 同步配置 fencing
第一段事务把 provider_id、provider.code、Provider 启用状态、账号启用状态、sync_enabled、credential_profile 和排序后的 region_allowlist 组成配置元组。网络发现结束后,第二段事务重新读取账号并计算同一元组;只要其中任一项变化,就抛 SYNC_CONFIGURATION_CHANGED。
配置比较通过后还会针对当前账号再执行一次 _validate_snapshot。因此发现期间修改白名单、凭据前缀或启用状态时,旧配置下取得的快照不会写入。它依赖数据库状态与二次检查,范围仅限当前同步协议;不要把它描述成通用分布式 fencing token。
15.5 Provider 禁用与账号禁用发生在创建 SyncRun 之前
锁住账号后,服务先检查 provider.is_active;为 False 时抛 PROVIDER_DISABLED。随后检查账号:is_active=False 或 sync_enabled=False 都抛 ACCOUNT_DISABLED。这两类预检失败不会创建 SyncRun,因为运行记录是在检查之后才创建。
Provider 网络失败则不同:它发生在 running SyncRun 已创建以后,所以该运行会被更新为 failed,并保存脱敏错误码与消息;已有资产不被改成 missing。
15.6 生命周期、stale 与 fencing 即时检查点
执行目录:devopsX/;说明:验证完整空快照把已有实例标记为 missing。
python manage.py test cmdb.tests.test_sync.SyncAccountTests.test_complete_empty_snapshot_marks_instances_missing
预期以 OK 结束。
执行目录:devopsX/;说明:验证恢复记录包含 lifecycle_state 与 missing_since。
python manage.py test cmdb.tests.test_sync.SyncAccountTests.test_restored_change_contains_lifecycle_fields
预期以 OK 结束。
执行目录:devopsX/;说明:验证过期执行被回收且不再阻塞。
python manage.py test cmdb.tests.test_sync.SyncAccountTests.test_stale_running_sync_is_failed_and_does_not_block_new_sync
预期以 OK 结束。
执行目录:devopsX/;说明:验证旧执行的快照被 fencing 拒绝。
python manage.py test cmdb.tests.test_sync.SyncAccountTests.test_superseded_sync_cannot_apply_its_snapshot
预期以 OK 结束;替代执行成功,旧执行抛 SYNC_RUN_SUPERSEDED。
执行目录:devopsX/;说明:验证发现期间配置变化会拒绝旧配置快照。
python manage.py test cmdb.tests.test_sync.SyncAccountTests.test_configuration_change_during_discovery_fences_snapshot
预期以 OK 结束;错误码为 SYNC_CONFIGURATION_CHANGED,资产数量仍为 0。
执行目录:devopsX/;说明:验证停用 Provider 在创建 SyncRun 前阻止同步。
python manage.py test cmdb.tests.test_sync.SyncAccountTests.test_disabled_provider_prevents_sync
预期以 OK 结束,并确认 SyncRun 数量为 0。
执行目录:devopsX/;说明:验证退役状态与审计记录处于同一原子事务。
python manage.py test cmdb.tests.test_sync.SyncAccountTests.test_retirement_rolls_back_when_audit_creation_fails
预期以 OK 结束;模拟变化记录写入失败后,实例仍为 present,retired_at 仍为空。
执行目录:devopsX/;说明:验证 retired 实例不会被自动恢复。
python manage.py test cmdb.tests.test_sync.SyncAccountTests.test_retired_instance_is_not_automatically_restored
预期以 OK 结束。
16 管理命令:初始化 Provider 与同步账号
16.1 bootstrap_cmdb 的最终文件
相对路径:devopsX/cmdb/management/commands/bootstrap_cmdb.py;内容:完整文件。
from django.core.management.base import BaseCommand
from cmdb.models import CloudProvider
class Command(BaseCommand):
help = "创建 CMDB v1 内置云厂商"
def handle(self, *args, **options):
providers = (
("fake", "Fake 教学云"),
("aliyun", "阿里云"),
)
for code, name in providers:
provider, created = CloudProvider.objects.update_or_create(
code=code,
defaults={"name": name},
)
action = "创建" if created else "更新"
self.stdout.write("%s云厂商:%s" % (action, provider))
命令维护内置 Provider 的 code 与 name。update_or_create 的 defaults 只有 name,没有 is_active,所以重复初始化不会把管理员已经停用的 Provider 擅自重新启用。测试 test_bootstrap_preserves_disabled_provider 锁定了这一行为。
16.2 sync_cloud_account 的最终文件
相对路径:devopsX/cmdb/management/commands/sync_cloud_account.py;内容:完整文件。
from django.core.management.base import BaseCommand, CommandError
from cmdb.models import CloudAccount, SyncRun
from cmdb.providers.base import ProviderError
from cmdb.services.sync import sync_account
class Command(BaseCommand):
help = "同步一个云账号的资产"
def add_arguments(self, parser):
parser.add_argument("account_key")
parser.add_argument(
"--provider",
dest="provider_code",
help="云厂商代码;同名账号存在于多个厂商时必须填写。",
)
parser.add_argument(
"--scenario",
default="default",
choices=["default", "empty", "partial", "failed"],
help="Fake Provider 教学场景",
)
def handle(self, *args, **options):
accounts = CloudAccount.objects.select_related("provider").filter(
account_key=options["account_key"]
)
if options["provider_code"]:
accounts = accounts.filter(provider__code=options["provider_code"])
matches = list(accounts[:2])
if not matches:
raise CommandError("云账号不存在:%s" % options["account_key"])
if len(matches) > 1:
raise CommandError(
"多个云厂商存在账号键 %s,请使用 --provider 指定云厂商。"
% options["account_key"]
)
account = matches[0]
try:
sync_run = sync_account(
account,
trigger=SyncRun.Trigger.COMMAND,
scenario=options["scenario"],
)
except ProviderError as exc:
raise CommandError("%s:%s" % (exc.code, exc.message)) from exc
if sync_run.status == SyncRun.Status.PARTIAL:
raise CommandError(
"同步部分完成(%s):%s"
% (sync_run.error_code, sync_run.error_message)
)
self.stdout.write(
self.style.SUCCESS(
"同步完成:状态=%s,发现=%s,新建=%s,更新=%s,未变化=%s,缺失=%s,恢复=%s"
% (
sync_run.status,
sync_run.discovered_count,
sync_run.created_count,
sync_run.updated_count,
sync_run.unchanged_count,
sync_run.missing_count,
sync_run.restored_count,
)
)
)
16.2.1 账号定位不能只假设 account_key 全局唯一
模型只保证同一 Provider 内 account_key 唯一,所以不同 Provider 可以都有 production。命令先按 account_key 查询,若提供 --provider 再缩小范围;发现两个匹配时强制要求指定 Provider。
16.2.2 场景参数只服务 Fake 教学
--scenario 的 choices 是 default、empty、partial、failed。真实 Provider 接收该参数也不会由 Fake 场景驱动,因为 adapter 分派按 provider.code 决定。
16.2.3 ProviderError 与 partial 都让命令返回失败
服务层抛出的 ProviderError 被转换为 CommandError,终端看到稳定 code 和脱敏 message,并返回非零退出状态。服务若正常返回 partial SyncRun,命令也主动抛 CommandError“同步部分完成”,避免自动化调度把部分结果当作完整成功;数据库中的 SyncRun 仍保留 partial 状态与错误详情。只有 succeeded 才输出成功统计。
16.3 从空数据库跑通 Fake 同步
下面步骤假定迁移已完成,并使用同一个 fake-demo 账号。命令逐条执行,不要把多条命令粘成一条。
16.3.1 初始化内置 Provider
执行目录:devopsX/;说明:创建或更新 fake 与 aliyun Provider。
python manage.py bootstrap_cmdb
首次执行会显示“创建云厂商”,重复执行会显示“更新云厂商”。两种结果都表明命令成功。
16.3.2 创建或更新 Fake 教学账号
执行目录:devopsX/;说明:准备无地域白名单的 Fake 账号。
python manage.py shell -c "from cmdb.models import CloudAccount, CloudProvider; provider = CloudProvider.objects.get(code='fake'); CloudAccount.objects.update_or_create(provider=provider, account_key='fake-demo', defaults={'name': 'Fake 演示账号', 'is_active': True, 'sync_enabled': True})"
该命令无业务输出即为正常;它不写任何凭据。
16.3.3 第一次 default:创建两台实例
执行目录:devopsX/;说明:执行完整默认快照。
python manage.py sync_cloud_account fake-demo --provider fake --scenario default
在此前没有实例时,预期状态为 succeeded,发现 2、新建 2、更新 0、未变化 0、缺失 0、恢复 0。
16.3.4 第二次 default:验证幂等
执行目录:devopsX/;说明:重复执行相同快照。
python manage.py sync_cloud_account fake-demo --provider fake --scenario default
预期发现 2、新建 0、更新 0、未变化 2,不新增实例变化记录。
16.3.5 empty:把范围内实例标记 missing
执行目录:devopsX/;说明:执行完整空快照。
python manage.py sync_cloud_account fake-demo --provider fake --scenario empty
预期状态为 succeeded,发现 0、缺失 2。实例仍在数据库和页面中,只是生命周期变为 missing。
16.3.6 partial:不能继续推断缺失
执行目录:devopsX/;说明:执行部分结果快照。
python manage.py sync_cloud_account fake-demo --provider fake --scenario partial
预期命令以非零状态结束,显示“同步部分完成”以及 PAGINATION_INCOMPLETE 的脱敏消息。数据库仍保存 status=partial、发现 0、缺失 0;前一步已经 missing 的实例保持原状态,不产生第二次缺失变化。
16.3.7 再次 default:恢复 missing 实例
执行目录:devopsX/;说明:让两台实例重新出现在快照中。
python manage.py sync_cloud_account fake-demo --provider fake --scenario default
预期恢复 2;两台实例回到 present,并各产生 RESTORED 变化。
16.3.8 failed:失败运行不改资产
执行目录:devopsX/;说明:触发可控 Provider 失败。
python manage.py sync_cloud_account fake-demo --provider fake --scenario failed
预期命令以非零状态结束,显示 PROVIDER_UNAVAILABLE 与脱敏消息。同步历史会新增 failed 运行,两台资产保持 present。
16.4 直接验证账号禁用分支
当前自动化测试对 Provider 禁用有专门用例;账号的 is_active 与 sync_enabled 共用同一个 ACCOUNT_DISABLED 分支。下面用可恢复的本地步骤验证 sync_enabled。
执行目录:devopsX/;说明:临时关闭账号同步开关。
python manage.py shell -c "from cmdb.models import CloudAccount; CloudAccount.objects.filter(account_key='fake-demo', provider__code='fake').update(sync_enabled=False)"
预期命令正常结束。
执行目录:devopsX/;说明:尝试同步已关闭同步开关的账号。
python manage.py sync_cloud_account fake-demo --provider fake --scenario default
预期命令失败并显示 ACCOUNT_DISABLED;因为预检发生在创建运行之前,所以同步历史不会新增本次 SyncRun。
执行目录:devopsX/;说明:恢复账号同步开关。
python manage.py shell -c "from cmdb.models import CloudAccount; CloudAccount.objects.filter(account_key='fake-demo', provider__code='fake').update(sync_enabled=True)"
预期命令正常结束。不要省掉恢复步骤。
16.5 管理命令行为测试
执行目录:devopsX/;说明:验证 bootstrap 不会重新启用已停用 Provider。
python manage.py test cmdb.tests.test_sync.SyncAccountTests.test_bootstrap_preserves_disabled_provider
预期以 OK 结束。
执行目录:devopsX/;说明:验证管理命令把 partial 报告为非零失败。
python manage.py test cmdb.tests.test_sync.SyncAccountTests.test_sync_command_reports_partial_as_failure
预期以 OK 结束;测试同时确认数据库中的 SyncRun 保持 partial。
执行目录:devopsX/;说明:验证同名账号必须用 --provider 消歧。
python manage.py test cmdb.tests.test_sync.SyncAccountTests.test_sync_command_requires_provider_when_account_key_is_ambiguous
预期以 OK 结束;测试随后指定 fake 并确认命令成功。
17 页面入口:账号同步、同步历史与实例变化
17.1 页面调用链与权限边界
本章页面链路是:账号详情 POST 到 account_sync;服务返回后重定向回账号详情;同步历史列表链接到公开 UUID 详情;实例详情展示标签和全部变化。所有写入口都要求 POST 与对应权限。
| 入口 | 权限 | 方法 | 核心行为 |
|---|---|---|---|
| 账号详情 | view_cloudaccount | GET | 按其他权限决定是否显示实例和 SyncRun。 |
| 立即同步 | sync_cloudaccount | POST | Fake 可选场景;非 Fake 强制 default。 |
| 同步历史列表/详情 | view_syncrun | GET | 只有同时具备 view_computeinstance 才展示变化中的实例信息。 |
| 实例详情 | view_computeinstance | GET | 展示 Provider/人工标签、生命周期和 before/after。 |
| 保存人工标签 | change_computeinstance | POST | 只写 source=manual。 |
| 标记退役 | retire_computeinstance | POST | 保留历史并写 RETIRED 变化。 |
17.2 表单中的人工标签与 Fake 场景
下面是最终表单文件中的两个连续原文片段。每个片段都完整包含相应类,没有用占位内容替代代码。
相对路径:devopsX/cmdb/forms.py;范围:第 69~82 行,最终文件连续原文。
class ManualTagForm(forms.Form):
key = forms.CharField(label="标签键", max_length=100)
value = forms.CharField(label="标签值", max_length=255, required=False)
def clean_key(self):
return self.cleaned_data["key"].strip()
def save(self, instance):
return ComputeInstanceTag.objects.update_or_create(
instance=instance,
source=ComputeInstanceTag.Source.MANUAL,
key=self.cleaned_data["key"],
defaults={"value": self.cleaned_data["value"].strip()},
)
clean_key 去掉标签键首尾空格;save 固定 source=manual,并以 instance + source + key 更新或创建,所以不会覆盖 Provider 标签。
相对路径:devopsX/cmdb/forms.py;范围:第 104~114 行,最终文件连续原文。
class SyncAccountForm(forms.Form):
scenario = forms.ChoiceField(
label="Fake 教学场景",
required=False,
choices=(
("default", "默认:发现两台实例"),
("empty", "完整空快照:演示缺失"),
("partial", "部分快照:不标记缺失"),
("failed", "Provider 失败"),
),
)
场景 choices 与管理命令一致。字段 required=False 是为了让非 Fake 页面不必提交该值;视图仍把空值归一为 default。
17.3 路由与视图的最终相关代码
路由片段连续覆盖账号详情、同步、实例标签、退役以及 SyncRun 页面。
相对路径:devopsX/cmdb/urls.py;范围:第 14~22 行,最终文件连续原文。
path("accounts/<int:pk>/", views.account_detail, name="account_detail"),
path("accounts/<int:pk>/sync/", views.account_sync, name="account_sync"),
path("instances/", views.instance_list, name="instance_list"),
path("instances/export/", views.instance_csv_export, name="instance_csv_export"),
path("instances/<int:pk>/", views.instance_detail, name="instance_detail"),
path("instances/<int:pk>/tags/", views.manual_tag_save, name="manual_tag_save"),
path("instances/<int:pk>/retire/", views.instance_retire, name="instance_retire"),
path("sync-runs/", views.sync_run_list, name="sync_run_list"),
path("sync-runs/<uuid:public_id>/", views.sync_run_detail, name="sync_run_detail"),
账号详情与同步视图如下。同步视图完整展示 POST 限制、表单验证、Fake 场景处理、服务调用与消息反馈。
相对路径:devopsX/cmdb/views.py;范围:第 138~209 行,最终文件连续原文。
@login_required
@permission_required("cmdb.view_cloudaccount", raise_exception=True)
def account_detail(request, pk):
account = get_object_or_404(
CloudAccount.objects.select_related("provider"),
pk=pk,
)
can_view_instances = request.user.has_perm("cmdb.view_computeinstance")
can_view_sync_runs = request.user.has_perm("cmdb.view_syncrun")
context = {
"account": account,
"regions": account.regions.prefetch_related("availability_zones"),
"instances": (
account.compute_instances.select_related(
"region",
"availability_zone",
)[:10]
if can_view_instances
else None
),
"sync_runs": account.sync_runs.all()[:10] if can_view_sync_runs else None,
"can_view_instances": can_view_instances,
"can_view_sync_runs": can_view_sync_runs,
"sync_form": SyncAccountForm(initial={"scenario": "default"}),
}
return render(request, "cmdb/account_detail.html", context)
@login_required
@permission_required("cmdb.sync_cloudaccount", raise_exception=True)
@require_POST
def account_sync(request, pk):
account = get_object_or_404(
CloudAccount.objects.select_related("provider"),
pk=pk,
)
form = SyncAccountForm(request.POST)
if not form.is_valid():
messages.error(request, "同步参数无效。")
return redirect("cmdb:account_detail", pk=account.pk)
scenario = form.cleaned_data["scenario"] or "default"
if account.provider.code != "fake":
scenario = "default"
try:
sync_run = sync_account(
account,
requested_by=request.user,
trigger=SyncRun.Trigger.MANUAL,
scenario=scenario,
)
except ProviderError as exc:
messages.error(request, "同步失败(%s):%s" % (exc.code, exc.message))
else:
if sync_run.status == SyncRun.Status.PARTIAL:
messages.warning(
request,
"同步部分完成(%s):%s"
% (sync_run.error_code, sync_run.error_message),
)
else:
messages.success(
request,
"同步完成:发现 %s,新建 %s,更新 %s,缺失 %s,恢复 %s。"
% (
sync_run.discovered_count,
sync_run.created_count,
sync_run.updated_count,
sync_run.missing_count,
sync_run.restored_count,
),
)
return redirect("cmdb:account_detail", pk=account.pk)
实例详情、人工标签和退役三个视图如下。
相对路径:devopsX/cmdb/views.py;范围:第 286~328 行,最终文件连续原文。
@login_required
@permission_required("cmdb.view_computeinstance", raise_exception=True)
def instance_detail(request, pk):
instance = get_object_or_404(
ComputeInstance.objects.select_related(
"account",
"account__provider",
"region",
"availability_zone",
).prefetch_related("tags", "changes__sync_run"),
pk=pk,
)
return render(
request,
"cmdb/instance_detail.html",
{"instance": instance, "tag_form": ManualTagForm()},
)
@login_required
@permission_required("cmdb.change_computeinstance", raise_exception=True)
@require_POST
def manual_tag_save(request, pk):
instance = get_object_or_404(ComputeInstance, pk=pk)
form = ManualTagForm(request.POST)
if form.is_valid():
form.save(instance)
messages.success(request, "人工标签已保存。")
else:
messages.error(request, "人工标签未保存,请检查输入。")
return redirect("cmdb:instance_detail", pk=instance.pk)
@login_required
@permission_required("cmdb.retire_computeinstance", raise_exception=True)
@require_POST
def instance_retire(request, pk):
instance = get_object_or_404(ComputeInstance, pk=pk)
if retire_instance(instance):
messages.success(request, "实例已标记为退役,历史记录仍然保留。")
else:
messages.info(request, "实例已经是退役状态。")
return redirect("cmdb:instance_detail", pk=instance.pk)
同步历史列表与详情视图如下。
相对路径:devopsX/cmdb/views.py;范围:第 331~362 行,最终文件连续原文。
@login_required
@permission_required("cmdb.view_syncrun", raise_exception=True)
def sync_run_list(request):
sync_runs = SyncRun.objects.select_related(
"account",
"account__provider",
"requested_by",
)
page_obj = Paginator(sync_runs, 30).get_page(request.GET.get("page"))
return render(request, "cmdb/sync_run_list.html", {"page_obj": page_obj})
@login_required
@permission_required("cmdb.view_syncrun", raise_exception=True)
def sync_run_detail(request, public_id):
can_view_instances = request.user.has_perm("cmdb.view_computeinstance")
sync_runs = SyncRun.objects.select_related(
"account",
"account__provider",
"requested_by",
)
if can_view_instances:
sync_runs = sync_runs.prefetch_related("changes__instance")
sync_run = get_object_or_404(sync_runs, public_id=public_id)
return render(
request,
"cmdb/sync_run_detail.html",
{
"sync_run": sync_run,
"can_view_instances": can_view_instances,
},
)
17.3.1 为什么视图仍然保持很薄
视图只做 HTTP 方法、权限、表单、消息和重定向;它不复制任何快照验证、事务或生命周期逻辑。这样命令行和页面得到完全相同的对账行为。
17.3.2 非 Fake Provider 不接受页面伪造场景
即使请求手工提交 scenario=empty,只要 provider.code 不是 fake,视图就把场景重置为 default。这避免教学参数影响真实 Provider 分派。
17.3.3 partial 使用警告消息而不是成功消息
服务返回 partial SyncRun 时,页面使用 messages.warning 展示 error_code 与 error_message;只有完整成功才显示 created、updated、missing、restored 统计的成功消息。这样部分结果不会在界面上被伪装成完整成功。
17.4 账号详情模板
相对路径:devopsX/cmdb/templates/cmdb/account_detail.html;内容:完整模板文件。
{% extends "base.html" %}
{% block title %}{{ account.name }} - CMDB{% endblock %}
{% block content %}
<section class="page-heading split-heading">
<div><p class="eyebrow">{{ account.provider.name }}</p><h1>{{ account.name }}</h1><p><code>{{ account.account_key }}</code></p></div>
{% if perms.cmdb.sync_cloudaccount %}
<form method="post" action="{% url 'cmdb:account_sync' account.pk %}" class="sync-form">
{% csrf_token %}
{% if account.provider.code == "fake" %}{{ sync_form.as_p }}{% endif %}
<button class="button" type="submit">立即同步</button>
</form>
{% endif %}
</section>
<section class="detail-grid">
<article class="panel"><h2>账号配置</h2><dl class="detail-list"><dt>云厂商</dt><dd>{{ account.provider.name }}</dd><dt>凭据前缀</dt><dd><code>{{ account.credential_profile|default:"未配置" }}</code></dd><dt>地域白名单</dt><dd>{{ account.region_allowlist|join:", "|default:"全部地域" }}</dd><dt>状态</dt><dd>{% if account.is_active %}已启用{% else %}已停用{% endif %}</dd><dt>最后成功同步</dt><dd>{{ account.last_successful_sync_at|date:"Y-m-d H:i:s"|default:"从未" }}</dd></dl></article>
<article class="panel"><h2>发现范围</h2><p>地域 {{ account.regions.count }} 个</p><p>可用区 {{ account.availability_zones.count }} 个</p>{% if can_view_instances %}<p>实例 {{ account.compute_instances.count }} 台</p>{% endif %}</article>
</section>
{% if can_view_instances %}
<section class="panel">
<div class="section-heading"><h2>实例</h2><a href="{% url 'cmdb:instance_list' %}?account={{ account.pk }}">查看全部</a></div>
<div class="table-wrap"><table><thead><tr><th>实例 ID</th><th>名称</th><th>地域</th><th>状态</th><th>生命周期</th></tr></thead><tbody>{% for instance in instances %}<tr><td><a href="{% url 'cmdb:instance_detail' instance.pk %}"><code>{{ instance.provider_resource_id }}</code></a></td><td>{{ instance.name }}</td><td>{{ instance.region.name }}</td><td>{{ instance.get_normalized_status_display }}</td><td>{{ instance.get_lifecycle_state_display }}</td></tr>{% empty %}<tr><td colspan="5" class="muted">尚未发现实例。</td></tr>{% endfor %}</tbody></table></div>
</section>
{% endif %}
{% if can_view_sync_runs %}
<section class="panel">
<div class="section-heading"><h2>同步历史</h2><a href="{% url 'cmdb:sync_run_list' %}">查看全部</a></div>
<div class="table-wrap"><table><thead><tr><th>执行 ID</th><th>状态</th><th>发现</th><th>新建</th><th>更新</th><th>缺失</th><th>时间</th></tr></thead><tbody>{% for sync_run in sync_runs %}<tr><td><a href="{% url 'cmdb:sync_run_detail' sync_run.public_id %}"><code>{{ sync_run.public_id }}</code></a></td><td><span class="badge badge-{{ sync_run.status }}">{{ sync_run.get_status_display }}</span></td><td>{{ sync_run.discovered_count }}</td><td>{{ sync_run.created_count }}</td><td>{{ sync_run.updated_count }}</td><td>{{ sync_run.missing_count }}</td><td>{{ sync_run.created_at|date:"Y-m-d H:i:s" }}</td></tr>{% empty %}<tr><td colspan="7" class="muted">尚无同步历史。</td></tr>{% endfor %}</tbody></table></div>
</section>
{% endif %}
{% endblock %}
模板根据 perms.cmdb.sync_cloudaccount 决定是否显示同步表单;场景下拉框只对 fake 账号显示。表单使用 POST 与 CSRF token。实例和同步历史两个区域分别受 view_computeinstance、view_syncrun 控制。
17.5 同步历史列表与详情模板
相对路径:devopsX/cmdb/templates/cmdb/sync_run_list.html;内容:完整模板文件。
{% extends "base.html" %}
{% block title %}同步历史 - CMDB{% endblock %}
{% block content %}
<section class="page-heading"><p class="eyebrow">Sync Run</p><h1>同步历史</h1><p>每一次发现都保留状态、统计和脱敏错误,不把运行过程伪装成异步任务。</p></section>
<section class="panel"><div class="table-wrap"><table><thead><tr><th>执行 ID</th><th>云账号</th><th>触发</th><th>状态</th><th>发现</th><th>新建</th><th>更新</th><th>未变化</th><th>缺失</th><th>恢复</th><th>时间</th></tr></thead><tbody>{% for sync_run in page_obj %}<tr><td><a href="{% url 'cmdb:sync_run_detail' sync_run.public_id %}"><code>{{ sync_run.public_id }}</code></a></td><td>{{ sync_run.account.name }}</td><td>{{ sync_run.get_trigger_display }}</td><td><span class="badge badge-{{ sync_run.status }}">{{ sync_run.get_status_display }}</span></td><td>{{ sync_run.discovered_count }}</td><td>{{ sync_run.created_count }}</td><td>{{ sync_run.updated_count }}</td><td>{{ sync_run.unchanged_count }}</td><td>{{ sync_run.missing_count }}</td><td>{{ sync_run.restored_count }}</td><td>{{ sync_run.created_at|date:"Y-m-d H:i:s" }}</td></tr>{% empty %}<tr><td colspan="11" class="muted">尚无同步记录。</td></tr>{% endfor %}</tbody></table></div>{% include "cmdb/_pagination.html" %}</section>
{% endblock %}
列表同时展示 discovered、created、updated、unchanged、missing 和 restored,避免只用“成功/失败”掩盖本次实际发生的事情。
相对路径:devopsX/cmdb/templates/cmdb/sync_run_detail.html;内容:完整模板文件。
{% extends "base.html" %}
{% block title %}同步 {{ sync_run.public_id }} - CMDB{% endblock %}
{% block content %}
<section class="page-heading"><p class="eyebrow">{{ sync_run.account.name }}</p><h1>同步执行详情</h1><p><code>{{ sync_run.public_id }}</code></p></section>
<section class="stats-grid"><article class="stat-card"><strong>{{ sync_run.discovered_count }}</strong><span>发现</span></article><article class="stat-card"><strong>{{ sync_run.created_count }}</strong><span>新建</span></article><article class="stat-card"><strong>{{ sync_run.updated_count }}</strong><span>更新</span></article></section>
<section class="detail-grid"><article class="panel"><h2>执行信息</h2><dl class="detail-list"><dt>状态</dt><dd><span class="badge badge-{{ sync_run.status }}">{{ sync_run.get_status_display }}</span></dd><dt>触发方式</dt><dd>{{ sync_run.get_trigger_display }}</dd><dt>请求用户</dt><dd>{{ sync_run.requested_by|default:"系统" }}</dd><dt>开始时间</dt><dd>{{ sync_run.started_at|date:"Y-m-d H:i:s" }}</dd><dt>结束时间</dt><dd>{{ sync_run.finished_at|date:"Y-m-d H:i:s"|default:"—" }}</dd><dt>错误代码</dt><dd><code>{{ sync_run.error_code|default:"—" }}</code></dd><dt>错误信息</dt><dd>{{ sync_run.error_message|default:"—" }}</dd></dl></article><article class="panel"><h2>统计</h2><dl class="detail-list"><dt>发现</dt><dd>{{ sync_run.discovered_count }}</dd><dt>新建</dt><dd>{{ sync_run.created_count }}</dd><dt>更新</dt><dd>{{ sync_run.updated_count }}</dd><dt>未变化</dt><dd>{{ sync_run.unchanged_count }}</dd><dt>缺失</dt><dd>{{ sync_run.missing_count }}</dd><dt>恢复</dt><dd>{{ sync_run.restored_count }}</dd><dt>退役</dt><dd>{{ sync_run.retired_count }}</dd></dl></article></section>
{% if can_view_instances %}<section class="panel"><h2>本次变化</h2><div class="table-wrap"><table><thead><tr><th>实例</th><th>动作</th><th>字段</th><th>时间</th></tr></thead><tbody>{% for change in sync_run.changes.all %}<tr><td><a href="{% url 'cmdb:instance_detail' change.instance.pk %}">{{ change.instance }}</a></td><td>{{ change.get_action_display }}</td><td>{{ change.changed_fields|join:", " }}</td><td>{{ change.created_at|date:"Y-m-d H:i:s" }}</td></tr>{% empty %}<tr><td colspan="4" class="muted">该次同步没有资产变化。</td></tr>{% endfor %}</tbody></table></div></section>{% endif %}
{% endblock %}
详情显示开始/结束时间、触发来源、请求用户、错误码和错误消息。只有 can_view_instances 为真时才渲染变化表,防止只拥有运行审计权限的用户看到实例标识。
17.6 实例详情模板:标签与变化历史
相对路径:devopsX/cmdb/templates/cmdb/instance_detail.html;内容:完整模板文件。
{% extends "base.html" %}
{% block title %}{{ instance.name|default:instance.provider_resource_id }} - CMDB{% endblock %}
{% block content %}
<section class="page-heading split-heading">
<div><p class="eyebrow">{{ instance.account.provider.name }} / {{ instance.account.name }}</p><h1>{{ instance.name|default:"未命名实例" }}</h1><p><code>{{ instance.provider_resource_id }}</code></p></div>
{% if perms.cmdb.retire_computeinstance and instance.lifecycle_state != "retired" %}<form method="post" action="{% url 'cmdb:instance_retire' instance.pk %}" onsubmit="return confirm('确认将该实例标记为退役吗?不会删除历史记录。');">{% csrf_token %}<button class="button button-danger" type="submit">标记退役</button></form>{% endif %}
</section>
<section class="detail-grid">
<article class="panel"><h2>云上属性</h2><dl class="detail-list"><dt>云账号</dt><dd>{% if perms.cmdb.view_cloudaccount %}<a href="{% url 'cmdb:account_detail' instance.account.pk %}">{{ instance.account.name }}</a>{% else %}{{ instance.account.name }}{% endif %}</dd><dt>地域</dt><dd>{{ instance.region.name }}({{ instance.region.provider_resource_id }})</dd><dt>可用区</dt><dd>{{ instance.availability_zone.name|default:"未提供" }}</dd><dt>实例规格</dt><dd>{{ instance.instance_type|default:"未提供" }}</dd><dt>计算资源</dt><dd>{{ instance.vcpu }} vCPU / {{ instance.memory_mb }} MiB</dd><dt>操作系统</dt><dd>{{ instance.os_name|default:"未提供" }}</dd><dt>厂商状态</dt><dd>{{ instance.provider_status|default:"未提供" }}</dd><dt>标准状态</dt><dd>{{ instance.get_normalized_status_display }}</dd></dl></article>
<article class="panel"><h2>发现与生命周期</h2><dl class="detail-list"><dt>生命周期</dt><dd>{{ instance.get_lifecycle_state_display }}</dd><dt>私网 IP</dt><dd>{{ instance.private_ips|join:", "|default:"无" }}</dd><dt>公网 IP</dt><dd>{{ instance.public_ips|join:", "|default:"无" }}</dd><dt>云上创建</dt><dd>{{ instance.cloud_created_at|date:"Y-m-d H:i:s"|default:"未提供" }}</dd><dt>首次发现</dt><dd>{{ instance.first_seen_at|date:"Y-m-d H:i:s" }}</dd><dt>最后发现</dt><dd>{{ instance.last_seen_at|date:"Y-m-d H:i:s" }}</dd><dt>开始缺失</dt><dd>{{ instance.missing_since|date:"Y-m-d H:i:s"|default:"—" }}</dd><dt>退役时间</dt><dd>{{ instance.retired_at|date:"Y-m-d H:i:s"|default:"—" }}</dd></dl></article>
</section>
<section class="detail-grid">
<article class="panel"><h2>标签</h2><div class="tag-list">{% for tag in instance.tags.all %}<span class="tag tag-{{ tag.source }}">{{ tag.key }}={{ tag.value }} <small>{{ tag.get_source_display }}</small></span>{% empty %}<span class="muted">暂无标签。</span>{% endfor %}</div>{% if perms.cmdb.change_computeinstance %}<form method="post" action="{% url 'cmdb:manual_tag_save' instance.pk %}" class="inline-fields">{% csrf_token %}{{ tag_form.as_p }}<button class="button" type="submit">保存人工标签</button></form>{% endif %}</article>
<article class="panel"><h2>变更历史</h2>{% for change in change_page %}<details><summary>{{ change.get_action_display }} · {{ change.created_at|date:"Y-m-d H:i:s" }}</summary><p>字段:{{ change.changed_fields|join:", "|default:"无字段变化" }}</p><div class="change-grid"><div><h3>变化前</h3><pre>{{ change.before_data }}</pre></div><div><h3>变化后</h3><pre>{{ change.after_data }}</pre></div></div></details>{% empty %}<p class="muted">暂无变更记录。</p>{% endfor %}{% if change_page.paginator.num_pages > 1 %}<nav class="pagination" aria-label="变更历史分页">{% if change_page.has_previous %}<a href="?change_page={{ change_page.previous_page_number }}">上一页</a>{% endif %}<span>第 {{ change_page.number }} / {{ change_page.paginator.num_pages }} 页</span>{% if change_page.has_next %}<a href="?change_page={{ change_page.next_page_number }}">下一页</a>{% endif %}</nav>{% endif %}</article>
</section>
{% endblock %}
标签用 source 对应的类名区分来源,并显示中文来源。人工标签表单只对 change_computeinstance 权限显示。退役按钮只对 retire_computeinstance 权限且非 retired 实例显示,并通过 POST 提交。
变化历史使用 details 原生元素按需展开;before_data 与 after_data 放进 pre,保留结构。这里显示的是数据库审计内容,不重新请求 Provider。
17.7 浏览器即时检查点
- 使用上一篇已经创建且有相应权限的账号登录,打开
http://127.0.0.1:8000/cmdb/accounts/,进入 Fake 演示账号。 - 账号详情应显示 Fake 教学场景下拉框和“立即同步”;选择 default 后提交,应看到成功消息、两台实例与一条新的同步历史。
- 打开
http://127.0.0.1:8000/cmdb/sync-runs/,应看到公开 UUID、状态、触发方式和统计;点击 UUID 进入详情。 - 同步详情在具备实例查看权限时显示两条 CREATED 变化;撤销该权限后,页面仍可显示运行信息,但不应显示“本次变化”和实例 ID。
- 进入一台实例详情,应看到 Provider 标签、生命周期、发现时间与 CREATED 变化;保存 owner 人工标签后,Provider 标签仍保留。
- 选择 empty 同步后刷新实例详情,生命周期应为“本次未发现”,变化新增“标记缺失”;再运行 default 后新增“恢复”。
- 点击“标记退役”后实例保留,生命周期变为“已退役”,变化历史新增“退役”;以后 default 同步不会自动恢复它。
若写入口用 GET 访问,预期响应为 405;若用户缺少权限,预期为 403 或页面隐藏相应数据,不应依靠仅隐藏按钮来代替服务端权限。
17.8 页面测试即时检查点
执行目录:devopsX/;说明:验证有同步权限的用户可从页面触发同步。
python manage.py test cmdb.tests.test_views.CmdbViewTests.test_user_with_sync_permission_can_start_sync
预期以 OK 结束,并确认 SyncRun.requested_by 是当前用户。
执行目录:devopsX/;说明:验证页面把 partial 显示为警告。
python manage.py test cmdb.tests.test_views.CmdbViewTests.test_partial_sync_displays_warning
预期以 OK 结束;响应包含“同步部分完成”,数据库运行状态为 partial。
执行目录:devopsX/;说明:验证仅有 SyncRun 查看权限时隐藏实例变化。
python manage.py test cmdb.tests.test_views.CmdbViewTests.test_sync_run_reader_does_not_see_instance_changes
预期以 OK 结束。
执行目录:devopsX/;说明:验证人工标签与退役只接受 POST。
python manage.py test cmdb.tests.test_views.CmdbViewTests.test_manual_tag_and_retire_are_post_only
预期以 OK 结束;GET 返回 405,合法 POST 返回重定向并完成写入。
18 测试矩阵、完整回归与本篇验收
18.1 同步服务的完整测试文件
同步服务是本篇最关键的边界,下面给出最终完整测试文件。它包含注入式 adapter、静态快照、可控失败、配置变更与替代执行,以及 33 个核心测试。
相对路径:devopsX/cmdb/tests/test_sync.py;内容:完整文件。
from dataclasses import replace
from datetime import timedelta
from io import StringIO
from unittest.mock import patch
from django.core.management import call_command
from django.core.management.base import CommandError
from django.test import TestCase
from django.utils import timezone
from cmdb.models import (
CloudAccount,
CloudProvider,
ComputeInstance,
ComputeInstanceChange,
ComputeInstanceTag,
SyncRun,
)
from cmdb.providers.base import (
DiscoveredAvailabilityZone,
DiscoveredComputeInstance,
DiscoveredRegion,
DiscoveredTag,
DiscoveryScope,
DiscoverySnapshot,
ProviderError,
)
from cmdb.providers.fake import FakeProvider
from cmdb.services.sync import retire_instance, sync_account
class StaticAdapter:
def __init__(self, snapshot):
self.snapshot = snapshot
def discover(self, account):
return self.snapshot
class FailingAdapter:
def discover(self, account):
raise ProviderError("ACCESS_DENIED", "凭据无权读取资产。")
class UnexpectedFailingAdapter:
def discover(self, account):
raise RuntimeError("provider response containing private details")
class ChangingConfigurationAdapter:
def discover(self, account):
CloudAccount.objects.filter(pk=account.pk).update(
region_allowlist=["cn-hangzhou"]
)
return FakeProvider.scenarios["default"]
class SupersedingAdapter:
def __init__(self):
self.replacement_run = None
def discover(self, account):
SyncRun.objects.filter(
account=account,
status=SyncRun.Status.RUNNING,
).update(started_at=timezone.now() - timedelta(minutes=31))
self.replacement_run = sync_account(account, adapter=FakeProvider())
return DiscoverySnapshot()
class SyncAccountTests(TestCase):
def setUp(self):
provider = CloudProvider.objects.create(code="fake", name="Fake 教学云")
self.account = CloudAccount.objects.create(
provider=provider,
account_key="fake-demo",
name="Fake 演示账号",
)
def test_first_sync_creates_two_instances_and_changes(self):
sync_run = sync_account(self.account, adapter=FakeProvider())
self.assertEqual(sync_run.status, SyncRun.Status.SUCCEEDED)
self.assertEqual(sync_run.discovered_count, 2)
self.assertEqual(sync_run.created_count, 2)
self.assertEqual(ComputeInstance.objects.count(), 2)
self.assertEqual(
ComputeInstanceChange.objects.filter(
action=ComputeInstanceChange.Action.CREATED
).count(),
2,
)
def test_identical_snapshot_is_idempotent(self):
sync_account(self.account, adapter=FakeProvider())
initial_change_count = ComputeInstanceChange.objects.count()
sync_run = sync_account(self.account, adapter=FakeProvider())
self.assertEqual(sync_run.created_count, 0)
self.assertEqual(sync_run.updated_count, 0)
self.assertEqual(sync_run.unchanged_count, 2)
self.assertEqual(ComputeInstanceChange.objects.count(), initial_change_count)
def test_changed_provider_field_creates_one_update(self):
default_snapshot = FakeProvider.scenarios["default"]
sync_account(self.account, adapter=StaticAdapter(default_snapshot))
first_instance = replace(default_snapshot.instances[0], name="订单服务-已更新")
changed_snapshot = replace(
default_snapshot,
instances=(first_instance, default_snapshot.instances[1]),
)
sync_run = sync_account(self.account, adapter=StaticAdapter(changed_snapshot))
self.assertEqual(sync_run.updated_count, 1)
instance = ComputeInstance.objects.get(provider_resource_id="i-fake-hz-001")
self.assertEqual(instance.name, "订单服务-已更新")
change = instance.changes.first()
self.assertEqual(change.action, ComputeInstanceChange.Action.UPDATED)
self.assertIn("name", change.changed_fields)
def test_complete_empty_snapshot_marks_instances_missing(self):
sync_account(self.account, adapter=FakeProvider())
sync_run = sync_account(
self.account,
adapter=StaticAdapter(DiscoverySnapshot()),
)
self.assertEqual(sync_run.missing_count, 2)
self.assertEqual(
ComputeInstance.objects.filter(
lifecycle_state=ComputeInstance.LifecycleState.MISSING
).count(),
2,
)
def test_repeated_empty_snapshot_does_not_duplicate_missing_changes(self):
sync_account(self.account, adapter=FakeProvider())
sync_account(self.account, adapter=StaticAdapter(DiscoverySnapshot()))
first_missing_changes = ComputeInstanceChange.objects.filter(
action=ComputeInstanceChange.Action.MARKED_MISSING
).count()
sync_run = sync_account(
self.account,
adapter=StaticAdapter(DiscoverySnapshot()),
)
self.assertEqual(sync_run.missing_count, 0)
self.assertEqual(
ComputeInstanceChange.objects.filter(
action=ComputeInstanceChange.Action.MARKED_MISSING
).count(),
first_missing_changes,
)
def test_restored_change_contains_lifecycle_fields(self):
sync_account(self.account, adapter=FakeProvider())
sync_account(self.account, adapter=StaticAdapter(DiscoverySnapshot()))
sync_account(self.account, adapter=FakeProvider())
change = ComputeInstanceChange.objects.filter(
action=ComputeInstanceChange.Action.RESTORED
).latest("created_at")
self.assertIn("lifecycle_state", change.changed_fields)
self.assertIn("missing_since", change.changed_fields)
self.assertEqual(change.after_data["lifecycle_state"], "present")
self.assertIsNone(change.after_data["missing_since"])
def test_instance_zone_must_belong_to_instance_region(self):
snapshot = DiscoverySnapshot(
regions=(
DiscoveredRegion("cn-hangzhou", "杭州"),
DiscoveredRegion("cn-shanghai", "上海"),
),
availability_zones=(
DiscoveredAvailabilityZone(
"cn-hangzhou-i",
"cn-hangzhou",
"杭州可用区 I",
),
),
instances=(
replace(
FakeProvider.scenarios["default"].instances[0],
region_provider_resource_id="cn-shanghai",
),
),
)
with self.assertRaisesMessage(ProviderError, "地域与可用区所属地域不一致"):
sync_account(self.account, adapter=StaticAdapter(snapshot))
self.assertEqual(ComputeInstance.objects.count(), 0)
def test_missing_instances_are_restored(self):
sync_account(self.account, adapter=FakeProvider())
sync_account(self.account, adapter=StaticAdapter(DiscoverySnapshot()))
sync_run = sync_account(self.account, adapter=FakeProvider())
self.assertEqual(sync_run.restored_count, 2)
self.assertFalse(
ComputeInstance.objects.exclude(
lifecycle_state=ComputeInstance.LifecycleState.PRESENT
).exists()
)
def test_partial_snapshot_never_marks_existing_instances_missing(self):
sync_account(self.account, adapter=FakeProvider())
partial_snapshot = FakeProvider.scenarios["partial"]
sync_run = sync_account(self.account, adapter=StaticAdapter(partial_snapshot))
self.assertEqual(sync_run.status, SyncRun.Status.PARTIAL)
self.assertEqual(sync_run.missing_count, 0)
self.assertFalse(
ComputeInstance.objects.filter(
lifecycle_state=ComputeInstance.LifecycleState.MISSING
).exists()
)
def test_provider_failure_records_failed_run_without_changing_assets(self):
sync_account(self.account, adapter=FakeProvider())
with self.assertRaises(ProviderError):
sync_account(self.account, adapter=FailingAdapter())
failed_run = SyncRun.objects.first()
self.assertEqual(failed_run.status, SyncRun.Status.FAILED)
self.assertEqual(failed_run.error_code, "ACCESS_DENIED")
self.assertEqual(ComputeInstance.objects.count(), 2)
self.assertFalse(
ComputeInstance.objects.filter(
lifecycle_state=ComputeInstance.LifecycleState.MISSING
).exists()
)
def test_unexpected_provider_failure_is_logged_and_sanitized(self):
with patch("cmdb.services.sync.logger.error") as log_error:
with self.assertRaises(ProviderError) as context:
sync_account(self.account, adapter=UnexpectedFailingAdapter())
self.assertEqual(context.exception.code, "UNEXPECTED_ERROR")
self.assertNotIn("private details", context.exception.message)
sync_run = SyncRun.objects.get()
log_error.assert_called_once_with(
"Unexpected CMDB synchronization failure type=%s account_id=%s sync_run_id=%s.",
"RuntimeError",
self.account.pk,
sync_run.pk,
)
self.assertEqual(sync_run.status, SyncRun.Status.FAILED)
self.assertEqual(sync_run.error_code, "UNEXPECTED_ERROR")
def test_duplicate_region_is_rejected_before_writing_snapshot(self):
snapshot = DiscoverySnapshot(
regions=(
DiscoveredRegion("cn-demo", "演示地域一"),
DiscoveredRegion("cn-demo", "演示地域二"),
)
)
with self.assertRaisesMessage(ProviderError, "发现重复地域 ID"):
sync_account(self.account, adapter=StaticAdapter(snapshot))
self.assertEqual(self.account.regions.count(), 0)
self.assertEqual(SyncRun.objects.first().status, SyncRun.Status.FAILED)
def test_provider_sync_does_not_delete_manual_tags(self):
default_snapshot = FakeProvider.scenarios["default"]
sync_account(self.account, adapter=StaticAdapter(default_snapshot))
instance = ComputeInstance.objects.get(provider_resource_id="i-fake-hz-001")
ComputeInstanceTag.objects.create(
instance=instance,
source=ComputeInstanceTag.Source.MANUAL,
key="owner",
value="local-team",
)
without_tags = replace(default_snapshot.instances[0], tags=())
changed_snapshot = replace(
default_snapshot,
instances=(without_tags, default_snapshot.instances[1]),
)
sync_account(self.account, adapter=StaticAdapter(changed_snapshot))
self.assertTrue(
instance.tags.filter(
source=ComputeInstanceTag.Source.MANUAL,
key="owner",
value="local-team",
).exists()
)
self.assertFalse(
instance.tags.filter(
source=ComputeInstanceTag.Source.PROVIDER,
key="department",
).exists()
)
def test_provider_tag_only_change_is_audited_as_update(self):
default_snapshot = FakeProvider.scenarios["default"]
sync_account(self.account, adapter=StaticAdapter(default_snapshot))
first_instance = replace(
default_snapshot.instances[0],
tags=(DiscoveredTag("department", "security"),),
)
changed_snapshot = replace(
default_snapshot,
instances=(first_instance, default_snapshot.instances[1]),
)
sync_run = sync_account(self.account, adapter=StaticAdapter(changed_snapshot))
self.assertEqual(sync_run.updated_count, 1)
self.assertEqual(sync_run.unchanged_count, 1)
change = ComputeInstanceChange.objects.filter(
action=ComputeInstanceChange.Action.UPDATED,
instance__provider_resource_id="i-fake-hz-001",
).latest("created_at")
self.assertIn("provider_tags", change.changed_fields)
self.assertEqual(
change.after_data["provider_tags"],
{"department": "security"},
)
def test_allowlist_scope_does_not_mark_excluded_region_missing(self):
default_snapshot = FakeProvider.scenarios["default"]
sync_account(self.account, adapter=StaticAdapter(default_snapshot))
self.account.region_allowlist = ["cn-shanghai"]
self.account.save(update_fields=["region_allowlist", "updated_at"])
shanghai_snapshot = replace(
default_snapshot,
regions=(default_snapshot.regions[1],),
availability_zones=(default_snapshot.availability_zones[1],),
instances=(default_snapshot.instances[1],),
scope=DiscoveryScope(
mode="allowlist",
region_ids=("cn-shanghai",),
),
)
sync_run = sync_account(self.account, adapter=StaticAdapter(shanghai_snapshot))
hangzhou_instance = ComputeInstance.objects.get(
provider_resource_id="i-fake-hz-001"
)
self.assertEqual(sync_run.missing_count, 0)
self.assertEqual(
hangzhou_instance.lifecycle_state,
ComputeInstance.LifecycleState.PRESENT,
)
def test_unscoped_account_rejects_allowlist_snapshot(self):
sync_account(self.account, adapter=FakeProvider())
snapshot = replace(
FakeProvider.scenarios["default"],
regions=(FakeProvider.scenarios["default"].regions[0],),
availability_zones=(FakeProvider.scenarios["default"].availability_zones[0],),
instances=(FakeProvider.scenarios["default"].instances[0],),
scope=DiscoveryScope(
mode="allowlist",
region_ids=("cn-hangzhou",),
),
)
with self.assertRaisesMessage(ProviderError, "未配置地域白名单"):
sync_account(self.account, adapter=StaticAdapter(snapshot))
self.assertEqual(
SyncRun.objects.latest("created_at").status,
SyncRun.Status.FAILED,
)
self.assertFalse(
ComputeInstance.objects.filter(
lifecycle_state=ComputeInstance.LifecycleState.MISSING
).exists()
)
def test_stale_running_sync_is_failed_and_does_not_block_new_sync(self):
stale_run = SyncRun.objects.create(
account=self.account,
status=SyncRun.Status.RUNNING,
trigger=SyncRun.Trigger.MANUAL,
started_at=timezone.now() - timedelta(minutes=31),
)
sync_run = sync_account(self.account, adapter=FakeProvider())
stale_run.refresh_from_db()
self.assertEqual(stale_run.status, SyncRun.Status.FAILED)
self.assertEqual(stale_run.error_code, "STALE_SYNC_RUN")
self.assertEqual(sync_run.status, SyncRun.Status.SUCCEEDED)
def test_superseded_sync_cannot_apply_its_snapshot(self):
adapter = SupersedingAdapter()
with self.assertRaises(ProviderError) as context:
sync_account(self.account, adapter=adapter)
self.assertEqual(context.exception.code, "SYNC_RUN_SUPERSEDED")
self.assertEqual(adapter.replacement_run.status, SyncRun.Status.SUCCEEDED)
self.assertEqual(ComputeInstance.objects.count(), 2)
self.assertFalse(
ComputeInstance.objects.exclude(
lifecycle_state=ComputeInstance.LifecycleState.PRESENT
).exists()
)
def test_configuration_change_during_discovery_fences_snapshot(self):
with self.assertRaises(ProviderError) as context:
sync_account(self.account, adapter=ChangingConfigurationAdapter())
self.assertEqual(context.exception.code, "SYNC_CONFIGURATION_CHANGED")
sync_run = SyncRun.objects.get()
self.assertEqual(sync_run.status, SyncRun.Status.FAILED)
self.assertEqual(sync_run.error_code, "SYNC_CONFIGURATION_CHANGED")
self.assertEqual(ComputeInstance.objects.count(), 0)
def test_disabled_provider_prevents_sync(self):
self.account.provider.is_active = False
self.account.provider.save(update_fields=["is_active", "updated_at"])
with self.assertRaises(ProviderError) as context:
sync_account(self.account, adapter=FakeProvider())
self.assertEqual(context.exception.code, "PROVIDER_DISABLED")
self.assertEqual(SyncRun.objects.count(), 0)
def test_fake_provider_filters_configured_region_allowlist(self):
self.account.region_allowlist = ["cn-hangzhou"]
self.account.save(update_fields=["region_allowlist", "updated_at"])
sync_run = sync_account(self.account, adapter=FakeProvider())
self.assertEqual(sync_run.status, SyncRun.Status.SUCCEEDED)
self.assertEqual(sync_run.discovered_count, 1)
self.assertEqual(
list(
ComputeInstance.objects.values_list(
"provider_resource_id",
flat=True,
)
),
["i-fake-hz-001"],
)
self.assertEqual(
list(
self.account.regions.values_list(
"provider_resource_id",
flat=True,
)
),
["cn-hangzhou"],
)
def test_fake_provider_rejects_unknown_allowlist_region(self):
self.account.region_allowlist = ["cn-unknown"]
self.account.save(update_fields=["region_allowlist", "updated_at"])
with self.assertRaises(ProviderError) as context:
sync_account(self.account, adapter=FakeProvider())
self.assertEqual(context.exception.code, "UNKNOWN_REGION_ALLOWLIST")
self.assertEqual(SyncRun.objects.first().status, SyncRun.Status.FAILED)
def test_unknown_fake_scenario_is_rejected(self):
with self.assertRaises(ProviderError) as context:
sync_account(self.account, adapter=FakeProvider("unknown"))
self.assertEqual(context.exception.code, "UNKNOWN_FAKE_SCENARIO")
def test_fresh_running_sync_still_blocks_new_sync(self):
SyncRun.objects.create(
account=self.account,
status=SyncRun.Status.RUNNING,
trigger=SyncRun.Trigger.MANUAL,
started_at=timezone.now(),
)
with self.assertRaises(ProviderError) as context:
sync_account(self.account, adapter=FakeProvider())
self.assertEqual(context.exception.code, "SYNC_ALREADY_RUNNING")
def test_empty_resource_identity_and_invalid_status_are_rejected(self):
empty_region = DiscoverySnapshot(
regions=(DiscoveredRegion("", "空地域"),),
)
invalid_status = DiscoverySnapshot(
regions=(DiscoveredRegion("cn-demo", "演示地域"),),
instances=(
DiscoveredComputeInstance(
provider_resource_id="i-demo",
region_provider_resource_id="cn-demo",
normalized_status="invalid",
),
),
)
with self.assertRaisesMessage(ProviderError, "空地域 ID"):
sync_account(self.account, adapter=StaticAdapter(empty_region))
with self.assertRaisesMessage(ProviderError, "标准状态无效"):
sync_account(self.account, adapter=StaticAdapter(invalid_status))
def test_partial_allowlist_snapshot_rejects_out_of_scope_region(self):
self.account.region_allowlist = ["cn-hangzhou"]
self.account.save(update_fields=["region_allowlist", "updated_at"])
snapshot = DiscoverySnapshot(
regions=(DiscoveredRegion("cn-beijing", "北京"),),
complete=False,
scope=DiscoveryScope(
mode="allowlist",
region_ids=("cn-hangzhou",),
),
)
with self.assertRaises(ProviderError) as context:
sync_account(self.account, adapter=StaticAdapter(snapshot))
self.assertEqual(context.exception.code, "OUT_OF_SCOPE_DISCOVERY_REGION")
self.assertFalse(self.account.regions.exists())
def test_complete_allowlist_snapshot_must_cover_its_scope(self):
self.account.region_allowlist = ["cn-hangzhou", "cn-shanghai"]
self.account.save(update_fields=["region_allowlist", "updated_at"])
snapshot = DiscoverySnapshot(
regions=(DiscoveredRegion("cn-hangzhou", "杭州"),),
scope=DiscoveryScope(
mode="allowlist",
region_ids=("cn-hangzhou", "cn-shanghai"),
),
)
with self.assertRaisesMessage(ProviderError, "没有覆盖地域白名单"):
sync_account(self.account, adapter=StaticAdapter(snapshot))
def test_bootstrap_preserves_disabled_provider(self):
provider = self.account.provider
provider.is_active = False
provider.save(update_fields=["is_active", "updated_at"])
call_command("bootstrap_cmdb", stdout=StringIO())
provider.refresh_from_db()
self.assertFalse(provider.is_active)
self.assertEqual(provider.name, "Fake 教学云")
def test_sync_command_reports_partial_as_failure(self):
with self.assertRaisesMessage(CommandError, "同步部分完成"):
call_command(
"sync_cloud_account",
"fake-demo",
provider_code="fake",
scenario="partial",
)
sync_run = SyncRun.objects.get()
self.assertEqual(sync_run.status, SyncRun.Status.PARTIAL)
self.assertEqual(sync_run.error_code, "PAGINATION_INCOMPLETE")
def test_sync_command_requires_provider_when_account_key_is_ambiguous(self):
aliyun = CloudProvider.objects.create(code="aliyun", name="阿里云")
CloudAccount.objects.create(
provider=aliyun,
account_key="fake-demo",
name="同名阿里云账号",
)
with self.assertRaisesMessage(CommandError, "--provider"):
call_command("sync_cloud_account", "fake-demo")
output = StringIO()
call_command(
"sync_cloud_account",
"fake-demo",
provider_code="fake",
stdout=output,
)
self.assertIn("succeeded", output.getvalue())
def test_retirement_rolls_back_when_audit_creation_fails(self):
sync_account(self.account, adapter=FakeProvider())
instance = ComputeInstance.objects.get(provider_resource_id="i-fake-hz-001")
with patch(
"cmdb.services.sync.ComputeInstanceChange.objects.create",
side_effect=RuntimeError("audit write failed"),
):
with self.assertRaises(RuntimeError):
retire_instance(instance)
instance.refresh_from_db()
self.assertEqual(
instance.lifecycle_state,
ComputeInstance.LifecycleState.PRESENT,
)
self.assertIsNone(instance.retired_at)
self.assertFalse(
instance.changes.filter(
action=ComputeInstanceChange.Action.RETIRED
).exists()
)
def test_retiring_missing_instance_clears_missing_state_in_audit(self):
sync_account(self.account, adapter=FakeProvider())
sync_account(self.account, adapter=StaticAdapter(DiscoverySnapshot()))
instance = ComputeInstance.objects.get(provider_resource_id="i-fake-hz-001")
retire_instance(instance)
instance.refresh_from_db()
change = instance.changes.filter(
action=ComputeInstanceChange.Action.RETIRED
).latest("created_at")
self.assertEqual(
instance.lifecycle_state,
ComputeInstance.LifecycleState.RETIRED,
)
self.assertIsNone(instance.missing_since)
self.assertIn("missing_since", change.changed_fields)
self.assertIsNotNone(change.before_data["missing_since"])
self.assertIsNone(change.after_data["missing_since"])
def test_retired_instance_is_not_automatically_restored(self):
sync_account(self.account, adapter=FakeProvider())
instance = ComputeInstance.objects.get(provider_resource_id="i-fake-hz-001")
retire_instance(instance)
sync_run = sync_account(self.account, adapter=FakeProvider())
instance.refresh_from_db()
self.assertEqual(sync_run.restored_count, 0)
self.assertEqual(
instance.lifecycle_state,
ComputeInstance.LifecycleState.RETIRED,
)
self.assertIsNotNone(instance.retired_at)
18.1.1 四个测试 adapter
StaticAdapter原样返回指定快照,用于精确构造边界。FailingAdapter抛 ACCESS_DENIED,验证失败运行与资产保护。ChangingConfigurationAdapter在 discover 期间改写账号白名单,验证旧配置快照被拒绝。SupersedingAdapter把原 run 变成 stale,再启动替代同步,最后返回旧快照,用于证明运行状态 fencing。
18.1.2 33 个同步测试逐项对应的保证
| 测试方法 | 保证 |
|---|---|
test_first_sync_creates_two_instances_and_changes | 首次发现创建两台实例与 CREATED 变化。 |
test_identical_snapshot_is_idempotent | 相同快照重复执行只计 unchanged。 |
test_changed_provider_field_creates_one_update | 受管字段变化形成 UPDATED。 |
test_complete_empty_snapshot_marks_instances_missing | 完整空快照可在范围内标记 missing。 |
test_repeated_empty_snapshot_does_not_duplicate_missing_changes | 重复空快照不重复写缺失变化。 |
test_restored_change_contains_lifecycle_fields | 恢复变化审计生命周期和 missing_since。 |
test_instance_zone_must_belong_to_instance_region | 实例地域和可用区关系必须一致。 |
test_missing_instances_are_restored | missing 资产重新出现时恢复 present。 |
test_partial_snapshot_never_marks_existing_instances_missing | 部分快照永不推断缺失。 |
test_provider_failure_records_failed_run_without_changing_assets | Provider 失败保留 failed run 且资产不变。 |
test_unexpected_provider_failure_is_logged_and_sanitized | 意外异常只记录类型和内部 ID,公开信息不泄露底层详情。 |
test_duplicate_region_is_rejected_before_writing_snapshot | 重复地域在写库前被拒绝。 |
test_provider_sync_does_not_delete_manual_tags | Provider 同步不删除人工标签。 |
test_provider_tag_only_change_is_audited_as_update | 仅 Provider 标签变化也写审计。 |
test_allowlist_scope_does_not_mark_excluded_region_missing | 白名单外实例不被误标 missing。 |
test_unscoped_account_rejects_allowlist_snapshot | 无白名单账号拒绝 allowlist 快照。 |
test_stale_running_sync_is_failed_and_does_not_block_new_sync | 陈旧运行被回收且不阻塞新同步。 |
test_superseded_sync_cannot_apply_its_snapshot | 被替代运行不能应用晚到快照。 |
test_configuration_change_during_discovery_fences_snapshot | 发现期间配置变化会 fence 旧快照。 |
test_disabled_provider_prevents_sync | 停用 Provider 阻止同步。 |
test_fake_provider_filters_configured_region_allowlist | Fake Provider 按地域白名单过滤。 |
test_fake_provider_rejects_unknown_allowlist_region | Fake Provider 拒绝未知地域。 |
test_unknown_fake_scenario_is_rejected | Fake Provider 拒绝未知场景。 |
test_fresh_running_sync_still_blocks_new_sync | 有效 running run 阻止同账号并发同步。 |
test_empty_resource_identity_and_invalid_status_are_rejected | 空身份和非法标准状态被拒绝。 |
test_partial_allowlist_snapshot_rejects_out_of_scope_region | partial allowlist 仍拒绝范围外地域。 |
test_complete_allowlist_snapshot_must_cover_its_scope | 完整白名单快照必须覆盖声明范围。 |
test_bootstrap_preserves_disabled_provider | 初始化命令保留 Provider 停用状态。 |
test_sync_command_reports_partial_as_failure | 管理命令以失败状态报告 partial。 |
test_sync_command_requires_provider_when_account_key_is_ambiguous | 跨 Provider 同名账号必须消歧。 |
test_retirement_rolls_back_when_audit_creation_fails | 退役审计失败时事务整体回滚。 |
test_retiring_missing_instance_clears_missing_state_in_audit | 退役 missing 实例会清空 missing_since,并在变化前后值中完整审计。 |
test_retired_instance_is_not_automatically_restored | retired 实例重现时不自动恢复;同步代码也不会静默清空 retired 行已有的 missing_since。 |
18.2 其他相关测试一个也不要漏看
同步服务不能脱离模型、页面和 Provider 合同单独判断。最终冻结源码的相关测试分布如下;展开每个文件可以看到 Django 实际发现的方法名。
devopsX/cmdb/tests/test_models.py:8 项
CmdbModelTests(8 项)
test_provider_and_account_string_values_are_readabletest_account_identity_is_unique_inside_providertest_same_account_key_can_exist_for_another_providertest_account_validates_credential_profile_and_region_allowlisttest_account_form_applies_region_allowlist_validatortest_zone_and_instance_accounts_must_match_their_relationstest_invalid_relation_ids_remain_validation_errorstest_tag_sources_are_independent
devopsX/cmdb/tests/test_views.py:19 项
CmdbViewTests(19 项)
test_anonymous_user_is_redirected_to_logintest_authenticated_user_sees_home_countstest_superuser_can_open_all_read_pagestest_synchronized_models_are_read_only_in_admintest_authenticated_user_without_permission_gets_403test_home_hides_account_create_without_account_view_permissiontest_provider_reader_does_not_see_account_countstest_sync_run_reader_does_not_see_instance_changestest_account_detail_hides_instances_and_sync_runs_without_permissionstest_authorized_get_on_sync_endpoint_gets_405test_sync_permission_without_view_permission_is_rejectedtest_user_with_sync_permission_can_start_synctest_instance_detail_does_not_link_account_without_account_view_permissiontest_instance_filter_searches_name_and_iptest_instance_filter_rejects_malformed_relation_ids_without_500test_partial_sync_displays_warningtest_instance_change_history_is_paginatedtest_topology_shows_instances_without_availability_zonetest_manual_tag_and_retire_are_post_only
devopsX/cmdb/tests/test_aliyun_provider.py:18 项
AliyunSdkResponseTests(4 项)
test_describe_regions_rejects_missing_collection_nodetest_describe_zones_rejects_missing_collection_nodetest_describe_instances_rejects_missing_list_or_totaltest_describe_instances_accepts_explicit_empty_page
AliyunMapperTests(5 项)
test_instance_mapper_normalizes_fieldstest_instance_mapper_maps_pending_to_startingtest_instance_mapper_rejects_invalid_nonempty_creation_timetest_instance_mapper_reads_vpc_private_addressestest_instance_mapper_reads_public_and_eip_addresses
AliyunProviderTests(9 项)
test_discover_supports_multiple_regions_and_pagestest_empty_region_collection_is_rejectedtest_empty_zone_collection_cannot_mark_existing_assets_missingtest_region_allowlist_limits_discoverytest_second_page_failure_returns_partial_snapshottest_invalid_instance_response_cannot_mark_existing_assets_missingtest_premature_empty_page_returns_partial_snapshottest_unknown_allowlist_region_is_rejectedtest_missing_environment_credentials_are_reported_without_secret
Aliyun 相关测试同时包含纯映射、真实 SDK 模型形状边界和数据库对账保护,但仍然不访问真实阿里云账号。页面测试还锁定 Admin 同步对象只读、实例详情无账号权限时不生成账号链接,以及首页新增账号入口必须同时满足新增和查看权限。
18.3 按模块执行测试
执行目录:devopsX/;说明:运行模型测试模块。
python manage.py test cmdb.tests.test_models
预期发现 8 个测试并以 OK 结束。
执行目录:devopsX/;说明:运行同步服务与管理命令测试模块。
python manage.py test cmdb.tests.test_sync
预期发现 33 个测试并以 OK 结束。
执行目录:devopsX/;说明:运行页面与权限测试模块。
python manage.py test cmdb.tests.test_views
预期发现 19 个测试并以 OK 结束。
执行目录:devopsX/;说明:运行使用模拟客户端的 Aliyun Provider 测试。
python manage.py test cmdb.tests.test_aliyun_provider
预期发现 18 个测试并以 OK 结束。这里仍然没有真实云认证。
18.4 执行完整回归
执行目录:devopsX/;说明:运行项目全部测试。
python manage.py test
当前最终源码在 SQLite 测试路径下发现 105 个测试,系统检查 0 个问题,并以 OK 结束。这个结果包含账号、API、CSV 等其他模块测试。
该结果不应被表述为“真实 Aliyun 已认证通过”或“MySQL 集成测试已通过”:Aliyun 相关用例使用模拟客户端;本次回归使用 DEBUG 下默认 SQLite。切换 MySQL 后仍需在目标环境单独执行迁移、测试并做并发验证。
18.5 本篇完成标准
- 能解释 adapter、DTO、snapshot、scope、partial result 与 reconciliation 的分工。
- 能逐条说出最新 scope 合同,并解释 Fake empty + allowlist 为何仍返回范围地域。
- 能说明 SyncRun 与 ComputeInstanceChange 分别审计什么。
- 能画出“短事务 A → 事务外网络 → 短事务 B”的 sync_account 流程。
- 能准确区分 MySQL 行锁与 SQLite 写锁缓解,不宣称分布式锁。
- 能说明完整与部分快照在 missing 推断上的根本差异。
- 能演示首次创建、重复幂等、missing、restored、retired、Provider 失败、stale recovery 与 fencing。
- 能说明 Provider 标签为何不会删除人工标签。
- 能从管理命令和页面查看每次运行及每台实例的变化历史。
- 能够运行本章相关测试与完整 105 项回归,并准确限定验证范围。
到这里,第 10~18 章结束。下一篇继续进入后续能力,不在本篇提前展开。

浙公网安备 33010602011771号