Theme
Ideas Session — 2026-08-01 15:04
Phase 1 — General Weather Agent Proposal
User request
用户希望在认真阅读 SIREN、Omni-Weather、Zephyrus、RadarQA、EarthLink、TerraBench、EWE、HVR-Met、WeatherSyn,以及医疗智能体 MIRA 和 frontier agent 设计 AgentCaster、MemEvolve 后,改进其通用气象 Agent proposal。目标是在任务定义、数据多样性、Agent 框架、持续学习、Agent memory 与其他可行创新点上显著区别于既有工作。用户还要求参考 /Users/leo/Library/Mobile Documents/com~apple~CloudDocs/科研/气象agent/feedback.md 中关于“现实工作流迁移 vs. Agent 新能力”、气象知识先验与 workflow 注入、底座能力保留、memory 作用的讨论,但不必全盘接受;最终方案可以用带文档证据的 grilling 方式压力测试。
Assistant response and working approach
我说明会先对齐本地分析、原始论文证据和 feedback,再提炼既有工作覆盖与真正空白,最后形成可执行 proposal,并以文档证据压力测试关键假设。采用的研究工作流包括:sci-brain:survey 组织文献版图,sci-brain:brainstorm-ideas 生成并筛选研究方向,sci-brain:idea-writer 在方向收敛后成稿,grilling 逐项解决需要用户决定的研究取舍。由于当前没有专用论文检索 MCP,将以本地材料为主、网页与官方论文页面作核验。
Initial evidence from feedback.md
现有讨论把 wrapper 与真正 Agent 的分界放在:工作流是否随证据动态形成;是否有成本敏感的主动证据获取;来源可靠性是否作为场景依赖潜变量学习;能否识别静默错误并偏离默认流程;是否具备停止、澄清和校准拒答;是否在新区域、新天气型、新工具、新故障和版本漂移下泛化。提出的统一抽象是:专家 workflow 作为可检索、可偏离、可验证、可更新的程序性先验;实时 evidence-conditioned policy 决定下一步;后验真值驱动受验证、可回滚的 skill/reliability/failure memory 更新。
Current question to investigate before asking the user
上述定位是否已被 SIREN、EarthLink、EWE、HVR-Met 或 TerraBench 部分覆盖;若有,proposal 的最小不可替代算法贡献应如何进一步收紧。
Evidence checkpoint — prior-art map and proposal audit
Prior-art map
| Work | What it already occupies | What remains open |
|---|---|---|
| Zephyrus | Executable weather world; dynamic code synthesis; reflective tool execution; 49 task types / 2,230 samples | Tool usefulness, conditional source health, cost-aware stopping, calibrated abstention, non-stationary adaptation |
| AgentCaster | Agent chooses HRRR maps and soundings under a sounding budget, then submits operational probability polygons | It does not learn a source-selection policy, model source reliability, or isolate meteorology from tool/format failures; current LLMs underuse soundings and overforecast |
| EWE | Knowledge-guided extreme-weather diagnosis, dual code/content auditors, short-term memory | Standard ReAct scaffold; vague memory; no reliable human baseline, source trust or operational degradation |
| HVR-Met | Hypothesis–verification–replanning, negative-result short-term memory, domain workflow library, atomic index/figure evaluation | Search for confirming anomalies is not falsifiable causal diagnosis; no competing-hypothesis, abstention, real-time imperfect-input or cost evaluation |
| SIREN | End-to-end event→forecast→impact→response benchmark; case RAG, skill rehearsal and task-specific modeling | Historical experience and warning chain are already occupied; leakage, prospective validity, cost matching and operational utility remain weak |
| TerraBench | Heterogeneous EO/grid/GIS/simulation tool environment; process, artifact, numeric-tolerance and provenance evaluation | Frozen ReAct system; no source-health belief, active value-of-information policy, online adaptation, or hidden source-state benchmark |
| EarthLink | Multi-plan climate research, code execution, diagnosis, expert-gated plan/script write-back | “Self-evolution” is external episodic memory accumulation; no unseen task-family proof, scientific verifier, conflict governance or memory rollback evidence |
| RadarQA | Domain forecast-quality attributes, radar sequence assessment, hybrid human/script/LLM annotation | Single SEVIR/VIL domain; assessment is not closed-loop evidence policy and does not improve forecast decisions |
| Omni-Weather | Shared radar generation + understanding backbone; structured weather rationales | No trustworthy agent control loop; rationales lack faithfulness, generalization and causal intervention evidence |
| WeatherSyn | Real forecast-discussion text aligned with weather fields; claim ontology and report generation | It is report synthesis, not evidence acquisition; future NWP fields are needed, salience is leaked through aspect prompts |
| MIRA | Controlled partially observable environment, typed action space, sequential evidence requests, process-level and human comparison | Strong transferable pattern: legality ≠ appropriateness; it lacks explicit cost/value-of-information optimization and real-world nonstationarity |
| MemEvolve | Memory mechanism itself as an executable search object; multi-objective architecture evolution | Search is small, related-domain only, lacks clean evolution/test isolation, risk objectives, true cross-round content migration and scientific safety |
Audit of the original proposal
The original proposal correctly identifies the strongest unoccupied center: quality-aware active evidence acquisition with stopping/abstention under heterogeneous source reliability, cost and latency. It also correctly separates deterministic weather constraints, workflow priors, real-time evidence, policy and verifier.
Its main weakness is not that the idea is too narrow, but that the document surrounds one strong algorithmic question with too many simultaneous claims: five task tracks, a universal source environment, QC, reliability learning, value of information, conflict diagnosis, abstention, multi-agent specialization, public/professional communication, active learning and continual adaptation. In the current literature this breadth risks being judged as system integration: every individual ingredient has nearby prior art, while no single mechanism is specified deeply enough to be the unavoidable reason for success.
The term “source reliability” also needs a sharper estimand. A radar, station, NWP model and ensemble do not measure the same object, so reliability cannot be a universal scalar. It should be a posterior over source health and task-conditional utility, conditioned on variable, scale, lead, regime, version and available QC; it should separately model data integrity, representativeness and predictive skill.
Three candidate centers
- Reliability-centered active evidence agent (recommended). Formalize a non-stationary meteorological evidence POMDP with latent source-health states. Learn a belief updater and cost-aware query/verify/stop/abstain policy. Make versioned reliability memory a posterior state rather than a vector-store feature. First paper: severe precipitation, narrow spatial/temporal contract, strong query-all/fixed-workflow/oracle baselines.
- Falsifiable weather-science agent. Extend HVR-Met from anomaly collection to explicit competing hypotheses, negative evidence, counterfactual tests and calibrated “unresolved” outcomes. Strong scientific-method novelty, but it pivots from operational forecasting to retrospective causal diagnosis.
- Lifelong weather-agent operating system. Treat workflows, source cards, verifiers and memory policies as typed/versioned programs; meta-evolve them under performance, cost, leakage, regression and stability constraints. Most frontier-agent flavored, but too large and risky as a first paper without a narrower task/environment contribution.
Recommended program structure: route 1 as the first defensible paper and shared core; route 2 as a later scientific-reasoning track; route 3 as the long-horizon general-agent program after a reliable sequential benchmark exists.
Decision checkpoint — primary research center
Assistant question
是否愿意把第一篇工作的唯一主线收紧为路线 1:非平稳来源可靠性下的主动气象证据获取与停止策略?推荐答案是同意;将可证伪的科学诊断和 typed self-evolving weather-agent OS 留作建立可靠 benchmark 之后的扩展。
User decision
用户同意该定位。
Consequence
后续 proposal 将以 latent source-health belief、task-conditional source utility、cost-aware query/QC/verify/stop/abstain policy 和 version-aware reliability memory 为不可替代主线。通用 LLM、workflow skills、tool schemas、verifier、procedural memory、多 agent 与多任务覆盖均降为环境、先验、实现或扩展,不再并列成为首篇论文的独立 headline。
Scope challenge — broad five-track benchmark versus focused first paper
Assistant recommendation before challenge
建议把首篇主任务固定为 CONUS 强降水 0–6 小时概率预报,按 0–1 h、1–3 h、3–6 h 分层,并把 6–24 h 留作迁移评测;理由是该时效已能形成雷达、站点、短临模型和 HRRR 之间真实的来源选择与冲突,同时保持可严格定义的真值、时间权限和成本。
User objection and proposed taxonomy
用户认为更广的任务覆盖可能更利于接收,并提出五类:短临预报、1–3 天短期预报、风险与政务决策支持、公众/行业服务生成,以及类似 Zephyrus 的数据读取、均值与绘图等 easy tasks。用户同时意识到工作量会很大。
Evidence-based response to prepare
Breadth helps a benchmark/resource paper when all tracks share one environment, one auditable state representation and one scientific hypothesis. It hurts a method paper when each track requires a different pipeline, ground truth and utility, because improvements cannot be attributed to the proposed mechanism. SIREN already covers event→forecast→impact→response; merely adding a fifth track is not a novelty claim. TerraBench also shows that broad executable coverage creates major artifact, code-version and evaluation burdens.
The five categories are best represented as task tracks of one shared agent, not five independent agents:
- Track 0: atomic grounding and deterministic scientific operations;
- Track 1: 0–2/6 h nowcasting;
- Track 2: 24–72 h short-range forecasting;
- Track 3: hazard→impact→action decision utility;
- Track 4: evidence-preserving communication transformation.
Ground-truth corrections are essential:
- Track 2 should use realized observations as primary GT; forecaster consensus is a human baseline or preference signal, not truth.
- Track 3 cannot use realized damage as an unconfounded direct label because mitigation, exposure and reporting affect damage; primary evaluation needs explicit hazard/exposure/action contracts and cost–loss or expert-reviewed decision utility.
- Track 4 should not use click-through rate as a primary success metric because sensational or overconfident messages may get more clicks; primary metrics should be evidence fidelity, salience, uncertainty, actionability and policy compliance.
- Track 0 is useful as an atomic diagnostic suite, but should not dominate the overall leaderboard or be called a separate easy agent.
Recommended compromise: preserve the five-track vision as a multi-paper platform, while making Tracks 0–2 the first paper's executable core and Tracks 3–4 smaller downstream transfer/consistency evaluations; alternatively, if all five must be equally mature, treat the effort as an 18–24 month benchmark/platform program rather than a single first submission.
Decision checkpoint — one integrated paper
User decision
用户要求先把五个 task tracks 与完整框架整合为一篇论文,再根据后续情况决定是否拆分;当前不以工作量为缩减依据,但最终 proposal 必须估算各部分的工作量分布。
Consequence
The proposal will be written as one benchmark-and-method paper with five tracks inside one executable environment. Workload will be estimated explicitly by work package, including data engineering, environment/tool contracts, benchmark/GT construction, agent method, continual-memory stream, expert annotation, evaluation, compute and release engineering. The paper still needs a single dominant scientific claim so task breadth functions as evidence of generality rather than five unrelated system demonstrations.
Venue decision — ICLR-first
User challenge
用户指出 Zephyrus 已被 ICLR 2026 接收,四位初审分数为 6/6/8/8,而其算法新颖性并不高,因此不认同把 ICLR bar 等同于强基础算法创新;用户决定按 ICLR 标准准备。
Updated interpretation
This objection is correct. Zephyrus demonstrates that ICLR can accept an application-driven agent paper whose main novelty lies in problem formulation, executable environment, benchmark, tool integration and evaluation rather than a new learning algorithm. Its acceptance case was strengthened by 49 task types / 2,230 samples, objective task-specific metrics, five backbones, tool ablations, direct versus reflective budget clarification, repeated-run variance, failure statistics, public code/data and a candid hard-task ceiling. The relevant lesson is not “algorithm novelty is mandatory,” but “the contribution type must be explicit and supported by unusually complete artifacts and causal ablations.”
Revised target shape
Prepare an ICLR-first system + benchmark + method paper:
- a new problem setting and five-track benchmark for non-stationary, heterogeneous, costly and conditionally reliable weather evidence;
- one shared general-purpose weather-agent environment with typed tools, evidence ledger, workflow priors, verifier and provenance;
- a reliability-aware active acquisition and continual-memory agent as the proposed system, without claiming a foundational new POMDP algorithm unless the implementation truly warrants it;
- extensive baselines, ablations, cross-backbone/OOD tests, cost accounting, human alignment and artifact release.
The differentiator from Zephyrus should be a new environmental difficulty and evaluation construct—source health, evidence sufficiency, operational degradation, delayed feedback and continual adaptation—not merely more tools or more task labels.
Benchmark organization decision and infrastructure ambition
Decision
用户同意采用“多数可规模化原子任务 + 少数完整事件链”的混合组织,但保留后续调整比例的权利。用户进一步提出:一旦发布,希望它成为气象 Agent 领域具有前瞻性和挑战性的基础设施,目标类比 ImageNet 对 image classification 的作用。
Consequence
The atomic/chain ratio must be treated as a versioned release configuration rather than the identity of the benchmark. Foundational infrastructure requires a stable underlying construct:
- a canonical event-centered episode schema;
- typed source/action/observation contracts;
- hidden source-health and corruption state;
- delayed outcome and feedback semantics;
- task-track extension protocol;
- versioned manifests, metrics and source adapters;
- public train/dev plus hidden and prospective test streams;
- an evaluation server that prevents benchmark contamination and supports rolling new events;
- reference baselines, human anchors and reproducible artifact bundles.
The benchmark should use a core + frontier design. The core maintains longitudinal comparability across releases; the frontier rotates new regions, hazards, tools, failures, task compositions and source versions. Our agent is a reference method, not the benchmark definition, so future algorithms can compete without inheriting our architecture.
Decision checkpoint — living benchmark and TerraBench-style paper shape
User decision
用户同意把 benchmark 设计为静态 Stable Core 加持续滚动的隐藏 Prospective Frontier。用户认可 benchmark、agent framework 和大规模 agents 横评同时构成贡献,并希望在 TerraBench 的“新数据集 + 自有框架 + 大横评”形态上进一步推进,从而提高 ICLR 接收概率。
Evidence-based qualification
This paper shape is plausible and well matched to ICLR's agent/AI4Science system-paper precedent. The claim of being “beyond TerraBench” must be operationalized as distinct constructs rather than greater size alone:
- latent and non-stationary source-health states;
- active query/QC/verify/stop/abstain actions;
- real and injected silent failures;
- delayed ground truth and verified continual memory;
- five-track weather-service chain;
- hidden prospective evaluation after release;
- paper–code–data metric consistency and frozen artifacts.
Breadth raises the release bar. A competitive infrastructure paper should match or exceed TerraBench/Zephyrus in reference baselines, cross-backbone coverage, cost normalization, repeat runs, human anchors, item-level artifacts, full metric implementation and code/data availability.
Decision checkpoint — hazard scope
User decision
用户同意 v1 采用“一主两辅”的灾种设计。
Frozen scope
- Primary hazard: severe precipitation / severe convection, with all five task tracks developed deeply;
- Transfer hazard A: high wind, testing different variables, spatial scales, source skill and decision-loss structure;
- Transfer hazard B: heat/cold extremes, testing slow evolution, weak radar dependence, climatological context and health/energy impacts.
The primary hazard supplies benchmark depth. The two transfer hazards test whether the source-health belief, active evidence policy, verifier and continual-memory mechanisms generalize beyond radar precipitation. Tropical cyclones, snowstorms, visibility, drought and other hazards remain prospective-frontier extensions rather than v1 requirements.
Decision checkpoint — source-health labels and annotation assistance
User decision
用户同意采用自然故障、可控反事实注入、隐藏组合故障和 prospective faults 的混合 source-health 设计,并提出必要时使用 GPT 辅助标注或从当地天气事件报告中提取标签与任务。
Annotation policy to carry forward
Use a provenance-tiered label stack:
- Tier A — objective truth: future observations, executable numerical oracles, official QC flags and paired corruption manifests;
- Tier B — documentary truth: NWS/NOAA/SPC event reports, forecast discussions, warnings, storm-event records and emergency/impact reports with exact source spans and timestamps;
- Tier C — LLM-assisted weak labels: structured event/claim extraction, candidate QA generation, failure taxonomy, report segmentation and rubric pre-scoring;
- Tier D — expert gold: meteorologist adjudication for ambiguous source conflicts, evidence sufficiency, decision reasonableness and communication quality.
LLM labels must retain model snapshot, prompt, source spans, confidence and disagreement metadata. They should not overwrite physical ground truth, and the same model family should not both generate and exclusively judge labels. High-impact hidden-test items need expert adjudication and reported inter-rater agreement.
Decision checkpoint — meteorological expert access
User decision
用户确认团队与气象专业人员有合作,可以在 proposal 中纳入专业人员标注与评测。
Consequence
Expert involvement will be designed as a core evaluation asset rather than an optional qualitative study. It should cover task/rubric validation, adjudication of ambiguous source conflicts and evidence sufficiency, blind review of decision and communication outputs, a same-information human baseline on a stratified subset, and calibration of LLM-assisted judges. The final workload estimate will separately budget expert hours and distinguish item creation, double annotation, adjudication and baseline execution.
Decision checkpoint — reference agent architecture
User decision
用户同意采用单一 Meta-Controller、按需 specialist、共享 evidence/belief ledger、独立 scientific verifier 和异步 verified-memory curator 的层级 Agent,而不是五个彼此独立、自由对话的 Agent。
Frozen architecture principle
The Meta-Controller owns the only authoritative task state, source-health belief, budget and stop/abstain decision. Observation/QC, forecast arbitration, impact/decision and communication specialists are invoked only when their expected marginal value is positive. Specialists exchange typed ledger entries rather than free-form facts. The communication specialist cannot introduce new meteorological claims. The verifier is logically independent from the generator, and long-term memory updates occur only after delayed truth or expert feedback. Single-agent ReAct and fixed/free multi-agent systems remain mandatory matched-budget baselines.
Decision checkpoint — typed memory-policy evolution
User decision
用户同意 v1 不仅更新记忆内容,还加入周期性的 typed memory-policy evolution;只允许离线或周期性演化,禁止单个事件后直接修改线上策略。
Frozen memory design
The long-term system includes structured reliability, procedural, failure, episodic and decision/user memories. Automatically evolvable objects are bounded memory-policy programs—write validation, source/regime/version partitioning, retrieval gates, ranking, consolidation, decay, expiration, conflict quarantine, rollback and schema migration. Candidate policies must pass type/property checks, temporally isolated selection/test evaluation, historical replay, regression and poison/staleness tests, cost/latency/storage constraints, and expert approval for high-risk changes. All releases are versioned and reversible.
Decision checkpoint — weather workflow as a soft policy prior
User decision
用户同意把 workflow 定义为可检索、可偏离、可验证、可演化的 soft policy prior;只有确定性的科学语义进入硬约束层。
Frozen knowledge-integration design
Hard scientific contracts encode units, time semantics, coordinates, variable types, physical ranges and permissions in typed APIs and verifiers. Procedural workflows are structured action graphs with applicability, mandatory checks, default actions, branches, failure modes, stop/abstain conditions, provenance and version. The foundation model may propose actions outside the retrieved workflow, skip low-value defaults or combine skills, but must log evidence-based deviations and pass independent verification. Required ablations include free ReAct, static SOP prompt, fixed expert pipeline, non-deviable retrieved workflow, deviable workflow prior, and the full reliability-aware policy.
Decision checkpoint — unified leaderboard
User decision
用户同意采用一个带严重错误门控的统一总分,同时保留五轨分数、可靠性诊断、预算分层和 Pareto 报告。
Frozen evaluation design
- The primary leaderboard uses a standard evidence/tool budget and reports a General Weather Agent Score (GWAS).
- Per-track utility is normalized against a floor and an oracle; the cross-track aggregate uses a geometric mean so that severe specialization cannot hide a collapsed track.
- Episode-level critical failures include temporal leakage, invalid unit/time/variable semantics, unsupported high-risk claims, ignoring a source already known to be broken, and unauthorized warning/action language.
- Critical-violation rate is reported separately and also gates the headline score; cost, latency, abstention calibration and evidence sufficiency remain visible rather than being collapsed away.
- Low-, standard- and high-budget tracks plus cost–quality Pareto curves distinguish active information acquisition from brute-force tool use.
- Static Core and hidden prospective Frontier results are reported separately.
Asset checkpoint — repository data/tool pipeline
A repository-wide search outside the paper-analysis notes found literature summaries and the evolving brainstorm log, but no executable GridRad/GHCNh/HRRR/GEFS/MRMS/NEXRAD data pipeline in /Users/leo/Documents/research-notes. The final proposal should therefore treat any already-developed data pipeline as an external team asset unless its location is later supplied, and should not claim that this repository contains an implementation. This makes the initial geographic/data scope a consequential workload decision.
Decision checkpoint — geographic scope
User decision
用户同意采用“CONUS 完整 Stable Core + 至少一个国际迁移 Prospective Frontier”的首版地域设计。
Frozen geographic design
- The CONUS Stable Core develops all five tracks and exploits the unusually reproducible alignment among public radar, satellite, surface observations, numerical guidance, official forecasts/warnings and event/impact reports.
- At least one geographically disjoint international Frontier evaluates transfer of the source-health ontology, active evidence policy, verifier and memory policy using globally reproducible or partner-authorized sources.
- The international Frontier need not duplicate the full five-track annotation density in v1; it should contain enough track and hazard diversity to test transfer rather than merely translating prompts.
- If partner access favors China, China can become the principal international Frontier while region-specific restricted data remain an optional expert evaluation tier rather than a prerequisite for reproducing the public Core.
- The infrastructure claim rests on the stable episode schema, source-health ontology, evaluation contracts and prospective service—not on claiming exhaustive country coverage in the first release.
Decision checkpoint — continual-learning boundary
User decision
用户同意 v1 冻结底座模型权重,把核心持续学习主张限定为经过验证的 typed memory 内容和 memory-policy 演化,并同时设置 Frozen-Agent 与 Continual-Agent 两条评测协议。
Frozen continual-evaluation protocol
- Frozen-Agent: resets long-term state between episodes; weights, workflow and policies remain fixed, while ordinary within-episode working memory is allowed.
- Continual-Agent: processes cases chronologically and may update long-term memory only after delayed observations, official reports or expert feedback become available; typed memory-policy changes occur only at scheduled offline checkpoints.
- Both protocols use the same frozen foundation-model weights and matched tool/data budgets so that improvement can be attributed to verified external memory rather than continued fine-tuning or extra inference.
- Evaluation obeys the causal order
event -> delayed truth -> validated update -> future event; final chronological blocks and prospective cases remain sealed. - Parameter-efficient adaptation or model self-training may be offered later as an explicitly separate Frontier track, not part of the first paper's central claim.
Decision checkpoint — challenge coverage versus deployment prevalence
User decision
用户同意采用高覆盖 Curated Challenge Set 与自然基率 Deployment Stream 的双采样设计。
Frozen sampling protocol
- Curated Challenge Set: deliberately stratifies severe events, near misses, benign/null cases, source conflicts, individual and compound failures, rare regimes and evidence-budget stress cases. It supports diagnostic coverage but is not presented as deployment prevalence.
- Deployment Stream: consists of contiguous time windows selected without conditioning on realized outcomes and preserves the natural event base rate. It measures false alarms, calibration, unnecessary queries, alert fatigue, stopping behavior and cost-sensitive utility.
- The two settings receive separate results and rankings; prevalence-sensitive metrics from the deployment stream are not inferred from the enriched challenge set.
- Every service track includes cases where the correct behavior is to avoid escalation: decline an unnecessary query, retain a low-risk forecast, recommend no costly intervention, avoid sensational communication, or abstain when evidence is insufficient.
- Full-chain episodes include false signals, weakening systems, observation artifacts and null events—not only memorable correctly forecast disasters.
Decision checkpoint — learned reliability-aware controller
User decision
用户同意把“LLM 推理 + 学习型 source-health/value/risk modules + 约束动作选择器”确定为 reference agent 的方法核心,而不是采用纯 prompt-based controller。
Frozen method boundary
- The frozen foundation model interprets tasks, forms meteorological hypotheses, proposes candidate evidence/actions and synthesizes outputs.
- Trainable, calibratable modules estimate the multidimensional source-health belief, expected marginal task utility of evidence and probability of critical error.
- A constrained selector chooses among querying, QC, cross-validation, specialist invocation, stopping and abstention by optimizing risk-adjusted value under explicit cost and latency budgets.
- Source health is represented by integrity, timeliness, spatiotemporal representativeness, regime-conditional skill and product/model version—not a single undifferentiated reliability score.
- Retrieved workflows supply action priors but cannot bypass semantic contracts or risk constraints.
- Paired counterfactual tests hold the weather situation fixed while modifying source availability/health, revealing whether the policy reacts to evidence state rather than replaying a fixed workflow.
- Pure prompting, free ReAct, static workflow and oracle-health/oracle-evidence variants remain necessary ablations and bounds.
Decision checkpoint — counterfactual evidence-subset replay
User decision
用户同意以完整历史事件档案上的反事实证据子集回放作为 controller 的主要训练机制,专家操作轨迹只作为辅助监督、校准材料与人类基线。
Frozen controller-supervision design
- Each historical event is reconstructed as a time-legal full-information bundle with delayed physical outcomes, documentary evidence and expert adjudications where needed.
- Training states are generated by masking sources, varying acquisition budgets/times and injecting naturalistic individual or compound faults.
- Candidate evidence is supervised through its measured marginal change in normalized downstream utility and critical-risk reduction, with query cost and latency retained explicitly.
- Objective physical/numerical ground truth supervises applicable tracks; downstream decision and communication labels use independent evaluators calibrated against expert judgments.
- Cross-fitting and evaluator separation prevent one model from generating, labeling and exclusively judging the same examples.
- Expert trajectories are not treated as a single correct workflow. They provide same-information human baselines, limited preference/calibration signals and qualitative error analysis.
- This protocol makes workflows a source of candidates and priors while learning evidence value from outcomes, and turns a single event archive into many paired source-health/budget counterfactuals.
Decision checkpoint — component and chain evaluation
User decision
用户同意五轨 benchmark 同时采用 Oracle-Upstream 与 Agent-Chain 两种模式。
Frozen cross-track evaluation design
- Oracle-Upstream Mode supplies a validated canonical upstream product and isolates the current track's capability.
- Agent-Chain Mode passes the same agent's typed upstream claims through the weather-service graph and measures propagation, recovery, uncertainty preservation and downstream restraint.
- Track 0 supports both Track 1 nowcasting and Track 2 short-term forecasting; either forecast branch can feed Track 3 decision support and Track 4 communication, while Track 3 may also feed Track 4.
- Ledger claims carry source provenance, issue/valid times, uncertainty and verification state. Downstream specialists cannot silently introduce or strengthen meteorological facts; a new claim requires new evidence and verification.
- Error decomposition distinguishes upstream forecast failure, impact/decision mapping failure, communication distortion, and successful downstream degradation or abstention in response to uncertain upstream claims.
- Both modes are reported; their eventual contribution to GWAS will be pilot-calibrated rather than fixed before item and human-difficulty distributions are known.
Decision checkpoint — bitemporal and versioned provenance
User decision
用户同意把 bitemporal/versioned provenance 设为正式评测数据的硬性准入条件,即使因此排除一部分无法重建历史可用时刻的数据。
Frozen temporal-integrity contract
- Each object records weather
valid_time, providerissue_time, benchmark-observedavailable_time, revision/product/model version, retrieval timestamp and immutable checksum. - Every tool call receives a
decision_timeand can return only the exact versions available by that point. - Final quality-controlled observations, post-event reports and expert reviews enter only through the delayed-feedback channel and cannot leak into the original episode.
- Sources whose historical availability or revision history cannot be reconstructed are marked
retrospective_only; they may support training, Track 0 or oracle analysis but are excluded from strict forecasting, Agent-Chain and prospective rankings. - The contract supports both leakage prevention and explicit modeling of source timeliness as a dimension of source health.
Decision checkpoint — continuous radar and satellite in the v1 Core
User decision
用户同意把连续 MRMS 与 GOES/GLM 从后续扩展提升为 v1 Core,并把 GridRad-Severe 明确限制在事件富集的 Curated Challenge Set。
Frozen initial source roles
- GridRad-Severe: high-value three-dimensional severe-convection/QC source for enriched challenge cases; it is not used to estimate natural event prevalence.
- Continuous MRMS or an equivalent continuous gridded radar archive: supports contiguous deployment windows, null events, false signals, near misses and prevalence-sensitive evaluation. Raw NEXRAD Level-II can remain a smaller specialist subset rather than an all-case storage requirement.
- GOES and GLM: become core multimodal nowcasting evidence instead of a later extension.
- GHCNh/ASOS-like surface observations: provide surface truth, precipitation/wind/temperature verification and natural station-failure cases.
- HRRR and GEFS: supply convection-allowing short-range guidance and ensemble uncertainty/context respectively.
- ERA5/WeatherBench 2: serve climatological context, historical analysis, training and interoperability baselines; reanalysis is not represented as contemporaneous raw observation.
- NWS/SPC forecast, warning and event documents plus static exposure layers: support Tracks 2–4 under the strict bitemporal availability contract.
- Distribution can use immutable event-cube manifests, checksums, preprocessing code and tool-mediated access rather than repackaging every raw radar/satellite object.
Decision checkpoint — Track 3 as structured decision making
User decision
用户同意把 Track 3 定义为结构化、可计算效用的决策选择任务,把政务专报等自然语言产物统一归入 Track 4。
Frozen Track 3/4 boundary
- Track 3 emits typed decision cards containing the hazard distribution, exposed assets, feasible actions, selected action, trigger/expiry, lead time, expected benefit, action cost, residual risk, supporting claims, uncertainty and escalation/abstention state.
- Evaluation combines hard legality/permission/time constraints, counterfactual cost–loss or resource-constrained utility, sensitivity to changed probabilities/exposure/action costs and blind expert review.
- Realized historical losses and actions are auxiliary evidence rather than unique ground truth because outcomes are confounded by defenses, reporting and chance.
- Track 4 converts validated forecast/risk/decision objects into public, industry, broadcast or government-facing communication. It may adapt audience, medium and readability but cannot silently change hazard severity, probability or recommended action.
- The separation makes Track 3 evaluate action selection under uncertainty and Track 4 evaluate faithful, useful communication instead of duplicating report-writing tasks.
Decision checkpoint — forecasting models as governed tools
User decision
用户同意把 Track 1/2 定位为 Agent 对已有预测模型的主动选择、融合、校准与可靠性管理,不把新的端到端天气预测模型训练纳入首篇论文。
Frozen predictive-model boundary
- Specialized extrapolation, NWP, ensemble and AI forecast systems are standardized tools with distinct validity, cost, latency, version and failure profiles.
- The agent decides whether and when to invoke them, diagnoses inputs/outputs, performs lightweight fusion and calibration, resolves disagreement and emits probabilistic hazard objects with calibrated uncertainty or abstention.
- Final numerical forecasts retain objective meteorological metrics such as CSI/POD/FAR and proper probabilistic scores; the method claim concerns reliable orchestration rather than a new weather backbone.
- A Fixed-Model Track gives every agent the same frozen forecast tools and is the primary policy comparison.
- A separately ranked Open-Model Track permits participant-provided forecasting models so that future model advances can enter the infrastructure without confounding the core agent leaderboard.
- The reference system may learn lightweight routing, fusion and calibration modules but does not train a new radar-generation or global-forecast foundation model.
Decision checkpoint — benchmark unit hierarchy and target scale
User decision
用户同意采用 Source Object、Event Cube、Decision Snapshot、Track Instance、Chain Episode 的分层单位,并以约 2,500–4,000 个独立事件窗、40,000–60,000 个分轨任务、600–1,000 个链式 episode 和 300–500 个双专家标注的高难链式 episode 作为 v1 目标档位。
Frozen scale and split principles
- Counts are reported at every hierarchy level; programmatic template paraphrases cannot masquerade as additional independent weather events.
- Roughly two-thirds of event cubes emphasize the primary severe-precipitation/convection hazard and one-third support high-wind and temperature-extreme transfer, with exact ratios pilot-adjustable.
- Track 0 supplies the largest automatically verifiable item pool without dominating GWAS; expensive chain/expert items are fewer but receive explicit coverage reporting.
- Deployment evaluation uses contiguous historical windows and aims to accumulate a full annual prospective cycle.
- Splits occur by storm system/event family and chronological block. Locations, times, tracks and templates derived from one weather process cannot cross train/test boundaries.
- Held-out task compositions, fault combinations, regimes and regions test compositional and source-health generalization.
- Final item counts remain target ranges until pilot difficulty, redundancy and statistical-power analyses are complete.
Decision checkpoint — autonomous and human-governed continual tracks
User decision
用户同意将 Autonomous Continual 设为可复现主榜,同时设置独立的 Human-Governed Continual 副榜。
Frozen continual-governance protocol
- Autonomous Continual: the submitted container, foundation weights and handwritten code are sealed before the hidden stream. Only declared typed-memory APIs and pre-registered policy-generation schemas may change persistent state.
- Automatic validation, promotion and rollback rules are fixed before evaluation; teams cannot inspect hidden cases and manually approve changes.
- Delayed observations, documents and any standardized expert-feedback objects are released identically and causally to all systems. Every persistent write and policy diff is logged.
- Human-Governed Continual: permits meteorologist review, quarantine, approval and rollback of high-risk memory/policy changes, while measuring expert interventions, minutes and incremental utility.
- The earlier expert-approval requirement applies to operational deployment and the human-governed track; the primary autonomous leaderboard instead uses pre-registered automated safety gates.
- Frozen-Agent, Autonomous Continual and Human-Governed Continual results remain separately identifiable.
Decision checkpoint — public evidence and restricted partner challenge
User decision
用户同意以公开可复现层作为主论文证据和主榜,把合作方受限数据定位为独立的高价值 Operational Challenge,不让核心 claim 依赖私有数据。
Frozen openness tiers
- Public Stable Core: primary paper claims and leaderboard results are reproducible from public sources, distributable derived objects or a universally accessible evaluation service, with manifests, bitemporal versions, preprocessing code, schemas and metrics released.
- Public Prospective Frontier: future cases and labels remain hidden during evaluation, but any team can submit through the same interface; retired cases are released when licensing permits.
- Partner Operational Challenge: restricted high-density observations, operational products, internal discussions or response feedback are evaluated through controlled infrastructure and reported separately as cross-region/institution external validity.
- No unique central-method conclusion may depend only on the restricted tier, and changes in a partner relationship must not disable the public benchmark.
- Expert annotations and rubrics derived from public cases should be released where consent/licensing permits; non-releasable judgments remain confined to the operational challenge with protocols and agreement statistics reported.
Decision checkpoint — falsification gates
User decision
用户同意把四组反证门槛正式写进 proposal,并用它们约束最终论文的贡献表述。
Frozen falsification logic
- Active reliability policy: under matched models, tools and budgets, the full method must improve task utility, critical-error rate and cost–quality tradeoffs over fixed workflows, free ReAct, all-source querying, active acquisition without source health and static historical-skill weighting.
- Generalization: gains cannot be confined to seen synthetic corruptions; clean performance, unseen compound faults, held-out regimes, at least one transfer hazard and prospective or cross-region performance are required.
- Continual memory: Autonomous Continual must improve future-event net utility over Frozen-Agent without increasing critical, stale-memory or negative-transfer failures. Historical replay alone cannot support the continual-learning claim.
- Benchmark validity: GWAS and track diagnostics must align interpretably with objective physical skill, decision utility and expert blind review, and rankings must not be arbitrarily unstable to small aggregation-weight changes.
Minimum effect sizes and confidence thresholds are selected on pilot/development data and frozen before hidden evaluation. Major comparisons use event-level bootstrap, repeated runs and transparent negative-result reporting. If a gate fails, the associated method or continual-learning claim is withdrawn even if the benchmark/data contribution remains publishable.
Final synthesis artifact
The agreed design has been synthesized into:
/Users/leo/Documents/research-notes/docs/projects/reliable-general-weather-agent-proposal-v2.md
The artifact includes the literature-gap matrix, formal problem, five-track benchmark, source/data/fault/annotation design, reference controller, workflow integration, typed continual memory, leaderboard, baseline and ablation grid, falsification gates, release lifecycle, safety boundaries, workload distribution and optional frontier innovations. It also incorporates K-MetBench and the Indian-monsoon decision-oriented benchmark as adjacent prior art discovered during the final audit.
The tentative acronym WEAVE was rejected because it is already used by multiple agent/vision benchmarks and tooling projects; the proposal therefore retains descriptive placeholder names pending a dedicated naming search.
Verification completed:
git diff --checkpassed;- VitePress production build passed with output redirected to a temporary directory;
- the only build notice was the repository's existing large-chunk warning;
- project index now links the new proposal.