# Research Loop · AI-Native 研究工作流模板 > Copyable artifact for《AI-Native Research》RES 13. The whole volume reduced to one runnable loop. > Isomorphic to engineering's spec-driven loop (Specify → Plan → Execute → Verify → Integrate → Learn), > with two research-specific load-bearing parts: **the knowledge graph comes first** (the evidence base > is the spec) and **integration takes priority over retrieval**. > > Evidence discipline: claims marked 〔evidence ledger〕 carry reliability; claims marked > 〔exploratory ledger〕 are leading indicators / boundaries / falsifiers, not proven facts. --- ## The six-step loop | # | Step | In one line | The load-bearing move | |---|------|-------------|-----------------------| | ① | FRAME | Stand up the evidence base; write the criteria | Build a traceable knowledge graph (legible to agents / traceable / falsifiable / integrable); write down explicit criteria for "worth believing / worth knowing." **The base is the spec — this precedes generation.** | | ② | GENERATE | Parallelize in-paradigm actions | Search / hypothesize / design experiments / analyze → hand to generation (RES 10 left cell). Each output carries evidence edges into the base. **What does not enter the base does not count.** | | ③ | TRIAGE | Believability ledger; suspend paradigm-level | Triage each batch on "evidence strength × paradigm distance": drop in-paradigm noise, integrate the believable, **suspend the paradigm-distant to seek discriminating evidence — do not kill as noise.** | | ④ | INTEGRATE | Synthesize across knowledge, not more retrieval | The human's scarce act: stitch never-juxtaposed claims into new understanding. Watch the ratio of integration artifacts to raw output. **Do not let generated-but-un-integrated knowledge pile into a mountain.** | | ⑤ | OWN VALUE | Set direction; leave the change-the-variable door | Give "worth" an owner; keep a human-initiated channel to "change the schema / change the variable," resisting the generation layer's conservative bias. | | ⑥ | FEED BACK | Errors become guardrails | Feed each retracted / refuted / mistakenly-killed novelty back as a new rule or node type in the base. Fewer mistakes next round. **This closes the loop.** | --- ## The believability ledger (③ in detail) — two axes, never one score Do **not** collapse credibility into a single number. A single score lets "distance from the existing-literature distribution" become the only proxy, which under-scores the genuinely novel. Book the two axes separately: | | **Evidence weak / untraceable** | **Evidence strong / traceable** | |-------------------------|---------------------------------|---------------------------------| | **Paradigm-distant (far)** | ⚠ **Do not kill yet — seek evidence.** A reframing is necessarily thin on evidence at birth (Einstein 1905, Lorentz both merely fit data at first). Suspend; go find the decisive discriminating evidence. | **Human — maybe a reframing.** Evidence holds while breaking the frame: send to a human, ask "did it switch to a new map?" | | **In-paradigm (near)** | **Doubt — in-paradigm noise.** Graph rules auto-downweight / block; no human bandwidth spent. | **Believe — integrate.** Traceable, base-consistent: auto-clear into the synthesis loop. | The cell to watch most is **weak × far**: deleting it as noise is exactly hypernormal science's killing move (RES 08). --- ## The in-paradigm / paradigm-level split (used in ② and ③) Operational test: **if you can write a machine-checkable acceptance criterion → left (hand to generation); if you can only appeal to "for whom, under which value frame" → right (stays with the human).** | Action | In-paradigm (→ generation) | Paradigm-level (→ human) | |--------|----------------------------|--------------------------| | Search | nearest-neighbor / RAG / citation tracing | — | | Hypothesize | next checkable gap where data is thickest | ask whether the variable set is the right level of description | | Design experiments | enumerate / sweep / ablate within a standard paradigm | cross-sensory analogy to embodied intuition | | Analyze | statistics / fitting / symbolic regression over predefined variables | judge whether novelty is a reframing or noise | | Judge | machine-checkable correctness | judge whether the conclusion is worth knowing | --- ## Starting path (do not build it all at once) 1. Pick **one** step where execution is already abundant but judgment is not yet externalized (literature synthesis, parameter sweeps). 2. Stand up a **minimal** traceable evidence base (① + ②). 3. Once it runs, add ③ the ledger and ④ integration. 4. **Reinvest the hours ② saves explicitly into ③④ — not into producing more papers.** (This is the single easiest place to go wrong; see RES 14 applicability boundary.) Three starts, one line: **stand up a traceable evidence base → use the believability ledger to choose where to inject human judgment → reinvest saved hours into integration and owning value.** --- ## Signals (are you running it right?) - **Right:** human-hours on in-paradigm actions keep falling; hours on paradigm-level judgment keep rising; share of left-cell outputs landing in the traceable chain on first pass rises; integration artifacts grow relative to raw output. - **Wrong (failure modes):** - forcing paradigm-level actions into the left cell ("let AI decide where to push the research") — handing value judgment away; - leaving in-paradigm actions in the right cell (humans still tracing citations by hand) — failing to abundify execution; - **hypernormal science** — output rises while topical breadth shrinks, "reframe" contributions stay low, citations concentrate (rising Gini): every metric green while science gets narrower. --- ## Source anchors - 〔evidence ledger · grade Ⅱ〕Hao, Xu, Li & Evans, "AI tools expand scientists' impact but contract science's focus," *Nature* 649(8099), 2026, DOI 10.1038/s41586-025-09922-y. ~41.3M papers; topical coverage −4.63%; scholar interaction −22%; citation Gini 0.754 vs 0.690. Observational bibliometrics; causal claims with care. - 〔exploratory ledger · grade Ⅳ–Ⅴ〕Djajadikerta, "Designing AI for Disruptive Science," *Asimov Press*, 2026-03-23, DOI 10.62211/29ej-27et. Opinion/review: hypernormal science; the map / Beck / Farr mechanism; meta-science's "model organism." An essay's argument, not data; its cited empirics each need tracing to original sources for grading.