Attoworld · September 2026 · Version 1

The AI-Native Research Playbook

How a researcher and a coding agent push a project from questions to findings and back, with the human's judgment where it counts and the machine's work everywhere else.

For theorists and experimentalists, students and group leaders. Two colours run through this page: red is what you own and do; blue is what the agent owns and does.

The bottleneck moved

It is worth being blunt about where the bottleneck went. The derivation, the convergence study, the literature search, the analysis pipeline, the first draft of the progress report: an agent produces each of them in minutes, to a standard that keeps rising. What has not sped up is you.

~250
words per minute you read
~40
words per minute you write
4
things working memory holds
1
of you, and you also run the experiment

So the scarce resource is bits through the human: what you must read, decide and specify. Machine work, tokens and compute are abundant. Every design choice in this playbook is settled by that one tie-breaker, and it inverts an intuition: the system should often do far more work in order to show you less. Run the limiting cases, the convergence test and an adversarial review, and report one conclusion with the checks beneath it. Reporting three candidate analyses for you to choose between is a defect.

It also fixes what the agent may ask you. You are asked what only you know: what is interesting, what was tried before, what the detector does above saturation, which of two framings matters. You are never asked what the agent could have found by reading the record. And it fixes what you read: after the first briefing, you read what changed, not the state of everything.

This is not an autonomous researcher. It is a system for delegating to agents what agents do well, getting from you what only you have, and removing the places where you were the bottleneck by accident rather than by design. You sign the paper and stand at the talk. Nothing enters the record and nothing leaves the project without your hand.

The loop

Research is a path from questions to plans to actions to findings, and back to questions. This is a more practical frame than "hypothesis, experiment, hypothesis": a question need not be a hypothesis, an action need not be an experiment, a finding need not be a confirmation. It is common to theory and experiment, to a doctoral student and to a group leader. Each station has a natural artifact, and each artifact appears as a byproduct of doing the work, never as an act of documentation.

The research loop: question, plan, actions, verify, findings, and back to question; communication leaves the loop. Question Plan Actions Verify Findings Communicate question document PLAN.md the work fresh context facts.md you send you state it, the agent writes it up the agent drafts, you approve the agent does; you do what only you can a second context reproduces every check yours, on your approval findings raise questions; decisions carry revisit conditions
Red stations are where your judgment enters; blue stations are where the machine works. Communication sits outside the loop because it leaves the project, and only you send.

Six stations, six artifacts

Six stations, each with an artifact that outlives the session and a point where your judgment enters. Four of them sit on the loop itself. Two are drawn apart from it: verification, because a result does not announce its own errors, and communication, because it leaves the project and only you send.

Station 1 · Question

A question with what would count as an answer

Key artifact
A question document: the question in one sentence, why it matters, what would count as an answer, what is already known and from where, constraints, priority.
You
State it in conversation. Answer four or five short prompts, as choices where possible. Endorse the answer criterion and the priority.
Agent
Interviews you, mines the record for what is already known, writes the document, proposes one line of it for the project page.
Gate
The framing is yours. The agent invites your position before offering its own, and never invents an answer criterion you did not endorse.

A question may be exploratory: "something is odd in the residual above 200 GW/cm²" is a legitimate question whose answer will be recognised rather than specified. The agent is forbidden to force precision on it. Curiosity that is not yet a question is still written down.

Station 2 · Plan

Approve plans, not outputs

Key artifact
PLAN.md: the questions it serves, the approach and why, ordered steps each with an owner and a completion criterion, which steps are independent, decision points, budget, risks, and the date you approved it.
You
Say how you would approach it, if you have a view. Read the plan. Change what you like. Approve.
Agent
Drafts the plan, marks which steps are yours and prepares what makes your part fastest: the email that asks the experimentalist for error bars, the parameter table, the two-option memo.
Gate
Work that outlasts a session runs on an approved plan, and so does work a session could finish whose cost you would want weighed first — restructuring code, a long run, a shared allocation. Otherwise a request your agent can finish in the session you asked it is just a request: it sequences its own moves, and where your harness already shows you that sequence, the system does not make you approve it twice.

Your judgment is cheapest before the work starts and most expensive after, when it has to be applied to outputs. The plan is where you spend it. It is also where the division of labour is written down: measure the 100 nm film and rerun the pipeline and inject a synthetic signal are both steps, with different owners.

Station 3 · Actions

The agent does; you do what only you can

Key artifacts
Code, notes, computed objects with provenance, and a finding document for every claim, with status tentative.
You
Turn the knob. Talk to the collaborator. Decide at the decision points. Correct in one sentence with the reason.
Agent
Derives, computes, analyses, searches. Runs every applicable check before you see anything. Writes the finding with its claim, what it means, its evidence, its provenance and its verification record.
Gate
Raw data is never touched. Every plot carries the script, the data, the code version, the parameters and the date. The agent may set no status but tentative.

Two things are worth naming here. Parallel sessions: independent plan steps run in separate contexts, and the plan says which steps are independent. Institutional knowledge as skills: when a procedure recurs, running this simulation pipeline, building this manuscript, the agent proposes a project-local skill, versioned with the work, and you approve it. This is how competence at a project's recurring work compounds instead of being rediscovered.

Station 4 · Verify

Every finding checks itself, and a stranger checks it again

Key artifact
The verification record inside the finding: every check applied, how it was run, what it produced; and the review record: who reviewed, in which context, when.
You
Read three lines: the claim, what it means, and which checks were not run. Ask for more only if you want it.
Agent
A second context that did not produce the finding reproduces every recorded check, runs the ones the producer omitted, attacks the claim, and sets the status: verified, refuted, or still tentative with what would settle it.
Gate
A check is reproducible or it is not a check: a command that was run, a computation that was evaluated, a location in a source that was read. "Checked, looks fine" does not qualify.

Nothing checks a research result automatically. A wrong derivation runs to the end and looks like a result; a wrong analysis produces a beautiful plot. So what stands in for a ground truth is named per kind of evidence: for a derived result, dimensional consistency, limiting cases, symmetries, a known special case, numerical evaluation against brute force, a second route derived rather than recalled; for a computed one, convergence in every discretisation parameter, a conserved quantity, a published figure reproduced, regeneration from a clean checkout; for a measured one, the pipeline rerun from raw data, a synthetic signal recovered, the expectation written down before the measurement compared with what came; for a literature claim, the sentence located in the source at the stated page, because a source that was not opened is a lead and not a citation. And for all of them, the review runs where the producer's blind spot cannot follow: in a fresh context.

Station 5 · Communicate

The agent drafts; you send

Key artifacts
Drafts: the email to the collaborator, the progress report, the talk outline, the manuscript section, the response to the referee.
You
Read it. Change it. Send it, submit it, sign it.
Agent
Writes for the recipient and the genre. States every claim at the strength its status allows and labels a tentative finding as such. Traces every number to a finding or a fact. Flags any commitment made in your name.
Gate
Nothing leaves the project by an agent's hand. Where the harness offers a deterministic gate, it enforces this; where it does not, the rule does, and you are told which.
Station 6 · Close the loop

The backlog fills itself; you choose from it

Key artifacts
The candidates note, where what was learned accrues; the pending note, where what waits on you accrues; and decisions.md, where every decision carries the condition under which to revisit it.
You
Say "sweep" at the end of a session that taught the system something. Approve what it proposes, tier by tier. Glance at what waits on you.
Agent
Notes every correction, preference and new fact as it happens. Notes every question a finding raises. Notices when a revisit condition has been met and says so. Proposes each item for the record, with a destination, and writes only on approval.
Gate
Nothing enters the record from conversation except through a sweep you approved. Nothing learns from a corrected output; things learn from a recorded reason.

This station could close on its own, with watchers that file new questions while you sleep. In this first version it deliberately does not. Literature monitors, instrument watchers and deadline trackers are named as extensions, to be added when a real failure justifies them. What the loop does close on its own is the part that matters most: a finding raises questions, a decision carries a revisit condition, and the next planning session puts both on the table.

Two territories

A project folder has two territories. The record is the agent's: six files, an inbox with two notes, and a filing cabinet, maintained under fixed rules. The work is yours: everything else, arranged however you like, including code in its own git repository and, if you want, a second agent maintaining it. Skills read the work freely and write to it when you ask, but never file, catalogue or prune it. The record points at the work; it never swallows it.

One line keeps the boundary honest: the filing cabinet is the agent's alone. INDEX.md promises that everything under it is indexed and reachable by a trigger, and a file you add, rename, move or delete there by hand breaks that promise silently — nothing checks it afterwards. Material goes into the inbox and the agent files it; when something should go, you approve a revision. Everywhere else you rearrange at will.

my-project/
├── AGENTS.mdpoints at the root convention and carries the standing rules; you own it
├── PROJECT.mdwhat and where: aim, state, status, next steps, open questions, exploration budget · 800 words
├── PLAN.mdhow: the current plan, with owners and completion criteria · you may edit it by hand
├── facts.mdwhat we rely on: each claim with provenance and a retirement condition
├── decisions.mdwhy: append-only, with revisit conditions
├── log.mdwhat happened: one line per operation, read forward from a date
├── INDEX.mdwhere to look, and when: a trigger map, never a summary
├── _WORKSPACE/in-tray and bench: what you drop in, and what the agent makes with no home yet
│ ├── candidates.mdwhat the agent learned, waiting for your approval at the next sweep
│ └── pending.mdwhat waits on you: queued decisions, and the agent's proposals
├── _FILED/the agent's filing cabinet; you do not edit this
└── research/yours, named by you: questions/, findings/, drafts/, explore/, model/ …

The six record files answer six different questions at six different times. Merging them into one README means reading all of them to answer any of them, which spends the only scarce resource. And the question and finding documents live in your territory on purpose: the record holds the claim of a verified finding with the document as provenance, but the document itself, with its evidence, method and checks, is yours to arrange by topic, by paper, by year, or however your field does it.

The agent speaks

A system that remembers, executes and verifies is an excellent assistant and still not a collaborator: it cannot disagree with you. The rule that it forms no view of what matters slides easily into forming no view of what is true, and then you are working alone with a very good machine. Four things prevent that, and they are the difference between a tool and a colleague.

A voice, and a place for it

When the agent disagrees with a plan, a claim or the record, or sees a conjecture, a better approach, a connection between two findings, or a question you have not asked, it says so in one line with its reason, at the point where it arises, and then does what you asked. The line goes into the pending note so it survives the session. You owe it a glance, nothing more. A proposal you pick up becomes a question; one you dismiss is removed at the next sweep. The agent forms these views because they are about what is true, possible or promising. It forms none about what you should want, and a proposal never carries a priority.

Exploration within a budget

Approved plans suit work whose shape is known. The frontier is also found by trying twenty things and keeping three, which agents do well and cheaply and you do not. So you may grant an exploration budget in the project page, say four hours of laptop compute per night in a folder set aside for it. Within it the agent explores without a plan, produces no findings, touches no record file, and reports only surprises. Anything interesting becomes a finding only after you turn it into a question and a plan step reproduces it with proper checks. No budget means no exploration.

Decisions that queue instead of stopping

When you are at the keyboard, a decision the agent needs is put to you and the work stops until you answer — through your harness's question tool where it has one, or by ending the turn and asking you in the conversation, which is the same act by another means. When you are not, in an overnight run, a subagent, or a session you told to continue without you, it is queued in the pending note and the agent continues with every step that does not depend on the answer. You return to a short list of what waited, not to a run that stopped at its first question. You are the system's judgment, never its blocking bottleneck.

Understanding as a deliverable

"Output is the conclusion, reasoning on request" is right for decisions and wrong for findings. A collaborator that leaves you unable to defend your own result at a talk has made you weaker while the record got stronger. So every finding carries what it means, in the words you would use, and accepting a claim into the record asserts that you could defend it. If you cannot, ask, and the agent explains as far as you want; it will not simply redo the work for you.

What the agent may not have a view on: what matters, what is interesting, what you should want, what to work on next. Those come entirely from what you have written down. A system that ranks your questions by its own lights has stopped being a mirror.

What changes

DimensionHow it is usually doneAI-native
Where a question livesIn your head, a notebook, a chat that is gone tomorrowA document with what would count as an answer, that an agent can work on
Where judgment is appliedLate, to outputs, after the work is doneEarly, to the plan; deviations update the plan
Who checks a result firstYou, reading everythingA second context that reproduces every check; you read three lines
How a result is trustedBy how well it was presentedBy its status, its verification record and its provenance
How the agent learnsIt doesn't; you repeat yourselfA correction with its reason goes into a file it reads every session
What you readEverything, to find the one thingWhat changed since you last looked; surprises before confirmations
What the agent says unaskedNothingDisagreement, conjecture, connection, one line each, where you will see it
While you are awayNothing happens, or it stops at the first questionDecisions queue; independent steps continue; exploration runs within its budget
What leaves the projectWhatever the chat window producedA draft with every claim at its status; you send
When the harness is replacedStart overThe files are the standard; give the spec to the new agent

Ten principles

If you take one thing from this page, take these ten. Everything else on it follows from them.

  1. Questions first. Every piece of research starts as a written question with what would count as an answer.
  2. Files, not chats. The interface between you and the agent is a set of files that carry over between sessions, agents and years.
  3. Approve plans, not outputs. Your judgment is cheapest before the work starts. The agent interviews you, writes the plan, you approve; deviations update the plan.
  4. Every finding checks itself. Limiting cases, convergence, a located citation, an adversarial review in a fresh context, before you read it. Every finding carries its status.
  5. Teach the system, not the chat. A correction with its reason goes into a file the agent reads every session. Say it once.
  6. Route judgment through the human, never data. Layered briefings, news not state, batched decisions as choices, questions only you can answer.
  7. Close the loop. Findings raise questions and decisions carry revisit conditions; the backlog fills itself and you choose from it.
  8. You send, you sign, you decide. Nothing leaves the project by an agent's hand.
  9. Provenance everywhere. Every plot, table and number traces to data, code, parameters and a date.
  10. The agent speaks. It disagrees, conjectures and proposes in one line where it arises, explores within a budget you grant, keeps working while you are away, and owes you an explanation of every finding. Never about what should matter; always about what is true or could be tried.

What survives the next model

Agents and models will be replaced, and then replaced again. The point of agreeing on this among colleagues is therefore not an implementation but a small set of standards simple enough to keep identical across people and across re-implementations:

Everything else is implementation, rebuilt freely: the wording of instruction files, the bodies of skills, hooks, subagents, scheduled runs, which agent runs it. A re-implementation is conformant if a project folder written under the old one is read correctly by the new one without conversion. Keep the standard, rebuild the rest, and no running project is ever disrupted by a better model.

Measuring it

If you want numbers, a few are cheap. Leading: time from a question stated to a plan approved; the fraction of findings verified at first review; how many decisions are waiting on you and for how long. Lagging: findings refuted after they entered the record; corrections you had to make twice, which mean the file was not read or was unclear. None of these is a target. A system built to lower a number can lower it by going quiet.

Start

You need a coding agent you already use, a folder where your projects live, and about an hour. You will not write code; your agent will.

  1. Read and edit the spec. It is written to be edited. If you change only one thing, change the lines that say what the agent asks you and what it settles alone.
  2. Give it to your agent with the one-paragraph build instruction in the getting-started guide. Your edited spec is what the agent builds from, and you keep it wherever you will find it again — change it later and ask for a rebuild; from it the agent writes a short convention document of what is local to you, the instruction files and the eleven skills for your harness, and tells you which rules it can enforce with a gate and which by instruction.
  3. Adopt one project. Not your most important one. Ask it what the state of the project is. If the answer is right, you are running.

Three skills carry the first week: ask, do, sweep. Four habits carry everything: drop things into the inbox, correct in one sentence with the reason, glance at what waits on you, and send everything yourself.

Everything this page describes lives in one repository: github.com/vsyakovlev77/ai-native-research. Beside this page you will find the spec, the principles, a walkthrough of one project from a question to a verified finding, a getting-started guide for Claude Code, Codex and Antigravity, and a FAQ. Take the spec, edit it, hand it to your agent. It is CC BY 4.0: edit it for your group and pass your version on.

Confidentiality. Unpublished data, a collaborator's unpublished result, grant text: decide per project what must not leave your machine and write it in one sentence in the project page. Every agent honours it. Members of the Max Planck Society have access to models hosted on GWDG's own infrastructure, which a project can name as the only model permitted for such material. For everyone: do not put into an external AI service anything you would not put into an email to a stranger.