One delegated task, all the way through
A researcher hands 380 abstracts to an assistant to screen against the inclusion criteria. Here is what PaperAlly does with that, step by step, and what each word means at the point where you need it.
Three words that do the work. Participation is what the assistant did. Reliance is how much the output was leaned on. Scholarly authority is who answers for the claim once it is in the paper. Most arguments about AI in research are three separate arguments wearing one word, and separating them is what makes the rest of this page possible.
The seven steps
R1 to R7 are the procedure. The rail on the left shows which rule you are in and which part of the record it fills.
- R1Task and output
Name the task, at the level a claim rests on
Not I used AI for my literature review. The unit is the operation whose output could change what the paper claims: decide, for each of 380 abstracts, whether it meets the inclusion criteria. One finite output, with a boundary around it.
Splitting the work this way is the whole trick. A review is not delegable or undelegable; the eleven different things you did inside it each have their own answer.
- R2Task and output
Say what the output is allowed to change
The screening decision decides which studies reach the evidence table, so it can change the finding. That is the authority at stake, and it is the reason the same operation is judged differently when it is only shortening a reading list.
PaperAlly types the task into a task family here. This one is Relevance screening, and the family carries a ceiling: the best warrant it can reach, whatever the output looks like on the day.
- R3Evidence access
Bind the evidence the system could actually see
What was the assistant given: the full text, the abstract alone, nothing at all? This is the question most disclosure statements never answer, and it is the one that decides whether the output can be checked by anyone.
Where a task asserts something about a source, evidence access becomes a binding rule: it can pull the warrant below the family ceiling, and no strong result anywhere else pulls it back up. An explicit no source was supplied is a valid answer and a much better one than silence.
- R4Reference standard, severe error
Fix what counts as right, and what counts as disqualifying
Before the run: what is the gold standard, who adjudicates a disagreement, and which single failure overrides a good average? For this family the engine already holds the answer.
Severe error: False exclusion: a relevant study silently dropped. A screening run at 92% agreement is not reassuring if the 8% is where the relevant trial went.
- R5Configuration
Run it, and log what shaped the output
Which model, which prompt, which settings, which tools, at what time. The run record is written as the run happens, which is the only moment it is cheap and the only moment it is accurate.
Configuration is the difference between a disclosure that can be repeated and one that has to be taken on trust.
- R6Named human check
Verify at the level the claim needs
Checking that a record exists is not checking that a field was read correctly, and neither is checking that a source supports the proposition it was cited for. The check has to sit at the level the claim rests on.
Here: a named person re-reads a sample of exclusions, because false exclusion is the severe error, and records the outcome. Accepted, corrected or rejected, with their name on it. There is no bulk mark-all-verified, on purpose.
- R7Named human check
Issue the warrant, and say when it reopens
The task closes with its warrant, the person who issued it, and the trigger that reopens it: a new model version, a change of protocol, a reviewer question.
For this family the standard's answer is V · delegable with verification. The rule attached to it: Use with protocol-specific sensitivity, false-exclusion review and documented human acceptance.
What the record now says
One task, six components, each of them a statement somebody else can check. This is the minimum warrant record, and it is what the warrant rests on.
| Task and output What finite output is delegated? | Decide inclusion for each of 380 abstracts. One output per abstract, in or out, with a reason code. |
| Configuration What shaped the output? | The assistant, its version, the prompt, the settings and the date, written by the run rather than remembered afterwards. |
| Evidence access What could the system inspect? | Title and abstract only. No full text was supplied, and the record says so in those words. |
| Reference standard What makes the output right or wrong? | The protocol's inclusion criteria, with disagreements adjudicated by the second reader. |
| Severe error and repeatability What error overrides average performance, and does the result repeat? | False exclusion is disqualifying. The run was repeated on a held-out sample to show the decisions were stable. |
| Named human check Who checks what, when, and with what outcome? | 20 of the 380 decisions re-read by a named second reader, weighted towards exclusions, outcome recorded per item. |
The three warrant states
Always written out in full, on every screen and in every export. The letters are for sorting; the words are what the reader needs.
The bounded output may enter the workflow after routine provenance and formatting checks.
Useful as a candidate but cannot support a claim until a named check is complete.
The evidence-bearing decision remains human.
Human-retained is not a telling-off. It is a statement about who holds authority for a decision. An assistant can still do the clerical work around an H task, propose candidates and tidy the prose. What it cannot do is be the thing the claim rests on.
What the reader in front of you gets
The tasks, their warrants, the six components and the sign-offs gather into one artefact. It is the same object every time. Only the name and the wording change, because a supervisor, an editor and a policy office are not asking the same question.
| Who is reading | What it is called, and what it emphasises |
|---|---|
| Your supervisor, or you | Your AI use record. Task by task, in plain words, with what you checked yourself. |
| A journal or a funder | AI disclosure record, containing the model attribution section, in the wording and the position that venue asks for. |
| A reviewer or editor | Review method note. What was used, for what, and on which parts of the report. |
| A policy or integrity office | Delegation record. The full set, with the rule that governed each decision. |
You decide who sees it and when. A record that is not sent has not been sent to anyone.
Three things this deliberately does not do
No likelihood score
There is no percentage, no probability and no AI-likelihood reading anywhere in the product, and there never will be. It reports which tasks were delegated, which rule applied, and who checked the output.
No verdict on your paper
The audit reports findings for a person to weigh. It does not grade the work, and a finding is a prompt for a conversation rather than a conclusion.
No proof of a negative
A clean record says the declared delegations were typed, warranted and checked. It says nothing about anything undeclared, because nothing in the system could know that.
Try it on something you have already written
Upload a finished manuscript and the deterministic audit reads it back to you: task families found, warrants assigned, R1 to R7, and the six components per task.