ARTICLE DETAIL

资讯详情

深耕编程入门与网站建设的一线实战洞察。

Context Engineering 的 /monitor.agent 实战:用模块化 System Prompt 构建可审计的 Agent 监控、告警与遥测管线

Context Engineering 的 /monitor.agent 实战:用模块化 System Prompt 构建可审计的 Agent 监控、告警与遥测管线 文档教程知识库人工智能提示工程【免费下载链接】Context-EngineeringContext engineering is the delicate art and science of filling the context window with just the right information for the next step. — Andrej Karpathy. A frontier, first-principles handbook inspired by Karpathy and 3Blue1Brown for moving beyond prompt engineering to the wider discipline of context design, orchestration, and optimization.项目地址https://gitcode.com/gh_mirrors/co/Context-Engineering点击查看免费下载导读/monitor.agent是 Context Engineering 仓库 AgenticOS 体系中面向系统/应用监控领域的标准 Agent 命令它以一份multimodal-markdown风格的模块化 System Prompt 为载体把「上下文/基础设施映射 → 指标选择 → 基线检查 → 持续采集 → 异常检测 → 告警 → 事件记录 → 审计反馈」的完整监控闭环固化为可复用的 Agent 行为。本文将以 .claude/commands/monitor.agent.md 为主体逐块拆解其 meta、instructions、context_schema、workflow、tools、recursion、examples 结构并结合仓库中的协议 Schema 与控制循环实现说明如何把这份提示词直接投入 Claude Code 等 Agent 运行时的 CLI 监控、健康检查、告警和遥测报告场景。一、定位AgenticOS 中的监控指挥中枢在 Context Engineering 的 .claude/commands/README.md 中仓库将一批*.agent.md命令组织为一个 Agentic Operating SystemAgenticOS每个 Agent 都以标准化结构实现特定领域的工作流。/monitor.agent在该体系中的职责是「System/service monitoring」——用于系统监控、告警、健康检查与遥测上报面向 Agent 与人类运维人员共用的 CLI 场景并强调**完全可审计auditable与自我改进self-improving**的工作流。它遵循与其他命令一致的模板骨架/monitor.agent.system.prompt.md ├── [meta] # 协议版本、审计、运行时、命名空间 ├── [instructions] # Agent 规则、调用方式、参数映射 ├── [ascii_diagrams] # 文件树、监控管线、告警/事件流转 ├── [context_schema] # JSON/YAML监控/会话/目标字段 ├── [workflow] # YAML监控各阶段 ├── [tools] # YAML/fractal.json工具注册表与管控 ├── [recursion] # Python反馈/事件循环 └── [examples] # Markdown样例仪表盘、告警日志这份「先定义元信息再写行为规则最后给 Schema、工作流、工具、循环与示例」的写法正是仓库强调的「Progressive Complexity渐进复杂度」范式在运维领域的落地从单条 promptatom逐步组装成可编排的监控 Agentorgan。二、协议头 [meta]版本、运行时与审计承诺/monitor.agent在文件开头用 JSON 声明了协议元信息{ agent_protocol_version: 2.0.0, prompt_style: multimodal-markdown, intended_runtime: [Anthropic Claude, OpenAI GPT-4o, Agentic System], schema_compatibility: [json, yaml, markdown, python, shell], namespaces: [project, user, team, infra, env], audit_log: true, last_updated: 2025-07-11, prompt_goal: Deliver modular, extensible, and auditable monitoring, health checking, alerting, and telemetry reporting—optimized for agent/human CLI and continuous improvement. }关键字段的工程含义字段值说明agent_protocol_version2.0.0Agent 提示词协议的语义化版本供运行时做兼容性检查prompt_stylemultimodal-markdown输出风格约定表格、ASCII 图、代码块、日志并存intended_runtimeClaude / GPT-4o / Agentic System目标运行环境保证跨厂商可移植schema_compatibilityjson/yaml/markdown/python/shell上下文 Schema 的载体格式族namespacesproject/user/team/infra/env变量命名空间用于区分来自项目、用户、团队、基础设施与环境的上下文audit_logtrue强制开启审计日志这是该 Agent 的核心设计承诺prompt_goal见上一句话说明该命令的职责边界从源码结构看.claude/commands/README.md 中所有命令共享同一套 meta 模板agent_protocol_version、prompt_style、intended_runtime、schema_compatibility、namespaces、audit_log、last_updated、prompt_goal/monitor.agent是其中把namespaces扩展为infra、env的典型——这正对应监控场景需要区分「基础设施对象」与「环境prod/staging/dev」的客观需求。三、行为准则 [instructions]Agent 的八条铁律[instructions] 区块用 Markdown 定义了/monitor.agent的行为契约You are a /monitor.agent. You: - Accept slash command arguments (e.g., /monitor targetapi metricslatency,uptime window1h alert95p) and file refs (file), plus shell/API output (!cmd). - Proceed phase by phase: context/infra mapping, metric selection, baseline check, continuous monitoring, anomaly detection, alerting, incident logging, feedback/audit loop. - Output clearly labeled, audit-ready markdown: metric dashboards, health summaries, anomaly logs, alert histories, incident timelines. - Explicitly declare tool access in [tools] per phase. - DO NOT skip baseline checks, alert configs, or audit logging. Do not alert without clear thresholds/context. - Surface all missed checks, alerting gaps, and false positives/negatives. - Visualize monitoring pipeline, alerting flow, and feedback/audit cycles. - Close with monitoring summary, audit/version log, unresolved incidents, and tuning recommendations.它同时定义了三种输入语法这是「Agent 与人类 CLI 共用」的关键接口约定命名参数targetapi metricslatency,uptime window1h alert95p其中alert95p表示第 95 百分位阈值percentile-based alerting其余为显式字符串文件引用file如contextinfra.md把文件内容注入上下文Shell/API 输出!cmd如context!kubectl get pods -A把命令输出作为监控输入。与 .claude/commands/README.md 中「!前缀执行 bash 命令、前缀引用文件」的通用约定完全一致。alert95p这种百分位写法也兼容 README 中/monitor serviceapi period24h alerttrue的布尔风格二者可视为同一参数在不同语境下的两种表达。行为约束中有三条「DO NOT / Surface」条款值得注意不允许跳过基线检查、告警配置与审计日志不允许在没有明确阈值/上下文的情况下告警必须暴露漏检、告警缺口与误报/漏报。这实质上把「告警疲劳alert fatigue治理」和「可解释告警」写进了提示词是监控 Agent 区别于简单「阈值触发器」提示词的核心差异。四、管线可视化 [ascii_diagrams]文件树与监控流水线该区块用 ASCII 图直观呈现两条关键结构。调用入口Slash Command 标准形态/monitor target... metrics... window... alert... contextfile ... │ ▼ [context/infra]→[metric_select]→[baseline]→[monitor/collect]→[anomaly_detect]→[alert]→[incident/log]→[audit/feedback] ↑_________________feedback/tuning/CI_________________|这条流水线揭示了/monitor.agent的编排本质所有监控工作被线性拆解为 8 个阶段且末端的audit/feedback通过一条反馈回路回到起点形成持续调优tuning与 CI 集成的闭环。它并非一次性执行脚本而是一个带反馈的稳态循环——这与仓库中 Control Loop 的哲学一脉相承。五、上下文 Schema [context_schema]监控会话的结构化输入/monitor.agent使用 YAML 定义了三层上下文结构保证无论谁调用输入字段都严格可解析monitor_context: target: string # 服务、应用、主机、集群等 metrics: [string] # latency, uptime, cpu, error, custom 等 window: string # 1h, 24h, rolling 等 alert: string # threshold, percentile, rule context: string provided_files: [string] constraints: [string] incidents: [string] args: { arbitrary: any } session: user: string goal: string priority_phases: [context, metric_select, baseline, monitor, anomaly, alert, incident, audit] special_instructions: string output_style: string team: - name: string role: string expertise: string preferred_output: string字段设计要点monitor_context是监控任务的「目标定义」target定义监控对象服务/应用/主机/集群metrics是度量列表延迟、可用性、CPU、错误率、自定义指标window定义时间窗口1h、24h、rolling 滚动窗口alert声明告警策略的类型阈值、百分位或规则constraints记录资源/权限/合规约束incidents挂接历史事件供基线对比args作为任意扩展参数的兜底字典。session描述会话语义priority_phases明确列出八阶段的执行优先级顺序使 Agent 在上下文受限时可据此裁剪output_style允许调用方指定输出风格。team是多人协作接口每个成员声明 role、expertise 与 preferred_output让监控结果可以按角色分发如 SRE 要时间线、管理层要摘要。这套 Schema 与 .claude/commands/README.md 描述的通用命令结构context_schema定义领域字段一致且比仓库中其他命令更强调team与session.priority_phases说明监控结果天然是多角色消费的。六、八阶段工作流 [workflow]每个阶段做什么、产出什么[workflow] 区块以 YAML 定义了 8 个监控阶段每个阶段都明确「做什么description」与「产出什么output」阶段职责输出context_infra_mapping解析 target、metrics、文件、窗口与约束澄清基础设施、目标与告警/事件需求上下文表、基础设施拓扑图、待澄清问题metric_selection选定/确认指标可用性、延迟、错误率等与相应窗口/阈值指标表、选择日志、规则矩阵baseline_check执行基线检查当前健康度、历史趋势、已知问题、告警配置基线仪表盘、历史表、告警配置日志monitor_collect持续采集指标、记录数据点、暴露事件监控仪表盘、指标日志、时间序列anomaly_detection用阈值、偏差或学习型规则检测异常标记漏报与误报异常日志、检测表、标记事件alerting触发、升级、记录告警暴露漏报/无效告警与告警疲劳风险告警日志、历史、通知表incident_logging记录事件、时间线与修复动作暴露未解决项与 RCA 触发点事件表、时间线、状态矩阵audit_feedback_loop审计所有阶段、工具调用、贡献者与版本检查点整合反馈与调优审计日志、版本历史、调优动作值得强调的三个设计决策基线检查前置baseline_check在持续采集之前执行为后续异常检测提供「正常值」参照系避免把常态波动误判为异常告警疲劳显式管理alerting阶段被要求「暴露漏报/无效告警与告警疲劳风险」把运维质量指标纳入 Agent 输出审计是阶段而非事后动作audit_feedback_loop是八阶段之一审计结果版本历史、调优动作会通过反馈回路影响下一轮监控这正是 [ascii_diagrams] 中那条回环线的语义。七、工具注册表 [tools]六个内部工具的协议化调用[workflow] 之后是 [tools] 区块声明了六个internal类型工具每个工具都带input_schema、output_schema、call协议与phases绑定工具 id职责调用协议绑定阶段infra_mapper映射目标/基础设施拓扑与服务依赖/infra.map{ targettarget, contextcontext }context_infra_mappingmetric_collector采集选定指标、摄入时间序列、快照健康度/metrics.collect{ targettarget, metricsmetrics, windowwindow }monitor_collectanomaly_detector用阈值/偏差/ML 规则检测指标异常/anomaly.detect{ logslogs, rulesrules }anomaly_detectionalert_manager管理告警触发、升级、通知与日志/alert.manage{ anomaliesanomalies, configconfig }alertingincident_logger记录、分类、时间线化事件供 RCA 与报告/incident.log{ alertsalerts, contextcontext }incident_loggingaudit_logger维护审计日志、指标事件与版本检查点/log.audit{ phase_logsphase_logs, argsargs }audit_feedback_loop这些工具的call字段采用/命名空间.动作{ 参数值 }的调用格式与仓库 60_protocols/schemas/protocolShell.v1.json 中定义的 Pareto-lang 操作模式^/[a-zA-Z0-9_]\.[a-zA-Z0-9_]\{.*\}$完全吻合。也就是说/monitor.agent的工具注册表并不是随意命名的而是与仓库协议层protocol shells的操作语法保持一致的——每个工具调用都可被协议层校验、记录与审计。各工具的输入输出 Schema 示例节选自文档- id: metric_collector type: internal description: Collect selected metrics, ingest time series, and snapshot health. input_schema: { target: string, metrics: list, window: string } output_schema: { logs: list, timeseries: dict } call: { protocol: /metrics.collect{ targettarget, metricsmetrics, windowwindow } } phases: [monitor_collect] examples: [{ input: {target: api, metrics: [latency,uptime], window: 1h}, output: {logs: [...], timeseries: {...}} }]工具与阶段一一对应除infra_mapper绑定context_infra_mapping外其余五个工具各绑定一个阶段形成「一个阶段一个工具」的清晰责任边界。这种显式声明也落实了 [instructions] 中「Explicitly declare tool access in [tools] per phase」的规则——Agent 在每一阶段能调用什么、不能调用什么都是可审计的。八、递归改进循环 [recursion]带审计的自我修正[recursion] 区块提供了监控 Agent 的循环骨架def monitor_agent_cycle(context, stateNone, audit_logNone, depth0, max_depth4): if state is None: state {} if audit_log is None: audit_log [] for phase in [ context_infra_mapping, metric_selection, baseline_check, monitor_collect, anomaly_detection, alerting, incident_logging ]: state[phase] run_phase(phase, context, state) if depth max_depth and needs_revision(state): revised_context, reason query_for_revision(context, state) audit_log.append({revision: phase, reason: reason, timestamp: get_time()}) return monitor_agent_cycle(revised_context, state, audit_log, depth 1, max_depth) else: state[audit_log] audit_log return state执行语义拆解顺序执行阶段对七个执行阶段逐一调用run_phase(phase, context, state)每阶段的结果写入state供后续阶段引用如anomaly_detection消费monitor_collect产出的日志与时间序列修订判断needs_revision(state)根据各阶段产出判断是否需要修订——例如基线偏差过大、告警配置缺口或漏检事件受控递归通过depth max_depth默认max_depth4限制递归深度防止无限循环修订原因与时间戳被写入audit_log审计随行audit_log全程贯穿最终挂入state[audit_log]随结果一起返回。该模式与仓库 20_templates/control_loop.py 中的ControlLoop类同构ControlLoop.run()同样维护iterations、results、context_manager在每次迭代中「生成响应 → 求值 → 记录 → 判断是否继续」并在达到max_iterations或stop_on_success时收敛。/monitor.agent的monitor_agent_cycle可以视为把这一通用控制循环实例化为监控领域阶段列表换成监控八阶段的产物max_depth4与ControlLoop的max_iterations5默认值一样都是「限制自我改进开销」的工程折中。九、端到端实战一条/monitor命令的完整生命周期[examples] 区块用一个api服务的例子演示了从调用到审计的完整输出。这里按原文档逐阶段还原并补充参数含义。9.1 Slash Command 调用/monitor targetapi metricslatency,uptime window1h alert95p contextinfra.mdtargetapi监控对象为 api 服务metricslatency,uptime采集延迟与可用性两个指标window1h分析窗口为 1 小时alert95p告警规则为第 95 百分位阈值contextinfra.md以前缀注入基础设施描述文件。9.2 Context/Infra Mapping阶段 1ArgValuetargetapimetricslatency,uptimewindow1halert95pcontextinfra.md9.3 Metric Selection阶段 2MetricWindowThresholdStatuslatency1h95p 300enableduptime24h99.9%enabled这里可见规则矩阵的精髓latency 用「95 分位小于 300ms」作为阈值uptime 用「24 小时窗口内 99.9%」作为可用性目标不同指标采用不同度量语义。9.4 Baseline Check阶段 3MetricValueHealthlatency122msgooduptime100%excellent基线阶段记录「当前值 健康评级」为后续异常判定提供参照。9.5 Monitoring/Collection阶段 4TimeMetricValueStatus16:10Zlatency111msok16:15Zuptime100%ok9.6 Anomaly Detection阶段 5TimeMetricValueAnomaly16:30Zlatency500msthreshold breachlatency 从基线 122ms 飙升至 500ms突破 95p 300ms 阈值被标记为「threshold breach」。9.7 Alerting阶段 6TimeAlertEscalatedRecipient16:30ZLatency spikeYesOn-call SRE告警被升级Escalated: Yes并定向通知 On-call SRE——体现team上下文对告警分发的指导作用。9.8 Incident Logging阶段 7IncidentTimeStatusRCA Triggerlatency50016:30Zresolvedyes9.9 Audit Log阶段 8PhaseChangeRationaleTimestampVersionAnomalyAdded thresholdAlert config2025-07-11 16:54Zv2.0IncidentLogged spikeRCA needed2025-07-11 16:55Zv2.0AuditVersion checkMonitoring loop2025-07-11 16:56Zv2.0审计日志以「阶段 变更 理由 时间戳 版本」的固定列结构记录每一次改动版本号随变更递增——这是整个 Agent 可审计性的最终落点也与 60_protocols/schemas/protocolShell.v1.json 中meta.version要求语义化版本^\d\.\d\.\d$的约定呼应。十、与 incident.agent / triage.agent 的协同从监控到事件的完整链/monitor.agent的incident_logging阶段产出的「事件表、时间线、RCA 触发点」恰好是仓库中另外两个运维 Agent 的输入20_templates/PROMPTS/triage.agent.md负责「intake → timeline → prioritize → investigate → evidence → root_cause → mitigate → audit」的应急分诊链路其incident上下文 Schemaid、type、summary、severity、detected_at、systems_affected、evidence_links可以直接消费/monitor.agent生成的告警与事件记录20_templates/PROMPTS/incident.agent.md负责完整的事故响应与根因分析RCA其intake_triage → timeline_construction → investigation → evidence_mapping → cause_effect_analysis → mitigation_planning → follow_up → audit_log八阶段正好承接监控 Agent 暴露的异常与事件。三者共享同一套模板骨架meta / instructions / ascii_diagrams / context_schema / workflow / tools / recursion / examples与同一套审计日志约定phase、change、rationale、timestamp、version 列结构因此可以无缝串成一条「持续监控 → 异常告警 → 事件分诊 → 根因分析 → 改进落地」的运维流水线。从源码结构看这正是 AgenticOS 设计者期望的编排方式每个 Agent 只负责一段清晰的职责边界通过标准化的输入输出 Schema 相互衔接。十一、部署与定制如何把/monitor.agent接入你的 CLI11.1 安装方式按照 .claude/commands/README.md 的约定命令有两种放置位置项目级.claude/commands/monitor.agent.md即当前仓库的位置只对当前项目生效个人级~/.claude/commands/monitor.agent.md对所有项目生效。在支持 slash command 的运行时如 Claude Code中放置后即可通过/monitor target... metrics... window... alert... contextfile ...触发。11.2 参数速查参数取值示例含义targetapi、db、prod-cluster监控对象服务/应用/主机/集群metricslatency,uptime,cpu,error逗号分隔的指标列表支持自定义指标window1h、24h、rolling时间窗口或滚动窗口alert95p、threshold、rule告警策略百分位/阈值/规则contextinfra.md、!kubectl get pods -A文件引用或命令输出注入args任意键值扩展参数的兜底入口11.3 定制建议替换工具实现六个internal工具infra_mapper、metric_collector、anomaly_detector、alert_manager、incident_logger、audit_logger的调用协议是标准化的可把metric_collector映射到 Prometheus/云监控 API把alert_manager映射到 PagerDuty/钉钉/邮件通知仅需保证输入输出 Schema 不变调整递归深度monitor_agent_cycle的max_depth4是自我改进循环的上限生产环境可结合告警疲劳指标误报率、漏报率适当调整扩充告警规则alert字段支持threshold、percentile、rule三种模式可在metric_selection的规则矩阵中按指标类型分别配置如延迟用百分位、可用性用目标值、错误率用计数阈值对接事件流水线将incident_logging输出接入 20_templates/PROMPTS/triage.agent.md 或 20_templates/PROMPTS/incident.agent.md实现监控告警到 RCA 的自动化衔接。十二、设计要点小结一份合格监控 Agent 提示词的五个特征回看/monitor.agent的完整结构可以提炼出可复用的设计原则模块化骨架meta / instructions / ascii_diagrams / context_schema / workflow / tools / recursion / examples 八个区块各司其职元信息与行为规则分离便于版本管理与跨运行时迁移Schema 先行monitor_context/session/team三层上下文结构让任何调用方人类或 Agent都能以确定性方式传参也为工具调用与审计记录提供了结构化基础阶段化编排 反馈闭环八阶段线性推进末端audit_feedback_loop通过递归循环回到起点实现基线调优与规则迭代协议化工具调用所有工具调用遵循protocolShell.v1的 Pareto-lang 语法60_protocols/schemas/protocolShell.v1.json可校验、可审计、可替换可审计性内建从audit_log: true的 meta 声明到audit_logger工具再到 examples 中带版本号的审计日志表审计贯穿 Agent 全生命周期。对希望构建「监控类 Agent 提示词」的开发者而言直接复用这份模板替换工具实现与指标规则即可获得一个自带基线检查、异常检测、告警升级、事件记录与审计闭环的监控 Agent——这正是 Context Engineering 仓库所倡导的「把工程实践固化为可复用上下文」的典型样例。赞分享文档教程知识库人工智能提示工程【免费下载链接】Context-EngineeringContext engineering is the delicate art and science of filling the context window with just the right information for the next step. — Andrej Karpathy. A frontier, first-principles handbook inspired by Karpathy and 3Blue1Brown for moving beyond prompt engineering to the wider discipline of context design, orchestration, and optimization.项目地址https://gitcode.com/gh_mirrors/co/Context-Engineering点击查看免费下载相关推荐用 Context-Engineering 的 /marketing.agent 构建模块化、可审计的营销战役智能体工作流用 Context Engineering 的 /marketing.agent 构建模块化、可审计的营销战役智能体工作流 导读本文以 Context Eng文档教程知识库人工智能提示工程Tambo CLI 深度指南面向 Agent 与 CI 的非交互式项目初始化和组件管理Tambo CLI 深度指南面向 Agent 与 CI 的非交互式项目初始化和组件管理 本文围绕 hydra ai 仓库中 Tambo CLI 的完整使用方式文档教程知识库人工智能提示工程PydanticAI Agent 开发 PRP 模板实战指南用 Context Engineering 构建可验证的最小可行 AgentPydanticAI Agent 开发 PRP 模板实战指南用 Context Engineering 构建可验证的最小可行 Agent 本篇技术指南围绕 c文档教程提示工程人工智能上一篇抖音内容批量下载3 步配置实现无水印视频与合集全量导出下一篇解锁AMD Ryzen性能密码SMUDebugTool深度调优指南创作声明:本文部分内容由AI辅助生成(AIGC),仅供参考
返回列表