I have spent enough time trying to reconstruct a paper from chat threads, tracked Word documents, and files called final_final_revised2 to be skeptical of another AI writing box. Most of them are fine for an email or a rough outline. They are not where I want to do the actual work of a manuscript.
A paper has more moving parts than prose. There is the text, yes, but also the references, tables, figure files, comments from coauthors, and the quiet little problems that only show up when the PDF finally builds. If an agent is going to help, I want it in the project with the rest of us, not producing an answer in a separate chat that I have to paste in and somehow reassemble later.
That is what I like about Sundial. It is a collaborative LaTeX editor built around the manuscript itself. The agent can propose a change, compile the paper, read the build log, and help check citations. The changes arrive as suggestions that can be kept or rejected. Sundial also supports live collaborators, two-way Overleaf sync, and local agents such as Claude Code and Codex joining the workspace.
None of that means a paper can run itself. It means the useful part of the work happens in one place, where a coauthor can see what changed and where it came from.
I Do Not Want One Agent Writing the Whole Paper
The bad version of AI coauthoring is easy to imagine: give three agents the draft and tell them to improve it. They all touch the introduction, each one adds a little more certainty to the conclusion, and you end up with polished prose that no one actually owns.
I would use Sundial differently. I would give an agent one narrow job, in a named part of the project, with a clear idea of what it should bring back. That might mean asking it to check a table, compare a small set of studies, tighten a methods section, or make sure the discussion does not claim more than the cited papers support.
A simple notes.md file is enough to make this concrete. The tags below are not a special Sundial routing feature. They are a convention for the team, which is the point. Anyone looking at the project can see what has been assigned, what is off limits, and what the agent is expected to return.
@evidence-agent
Work only in sections/background.tex, paragraphs 3 to 5.
Compare the two explanations for post-treatment imaging change. Read primary papers, not only abstracts. Keep the current argument unless the sources make it wrong. Propose at most 250 new words.
Bring back: a suggested diff, a source note with a DOI or URL behind every new factual claim, and a short list of what still needs author judgment.
Do not touch the abstract, methods, results, discussion, or unrelated bibliography entries.
That request is much more useful than “improve the background.” It gives the agent a boundary and gives the reviewer a checklist. It also stops the next agent from wandering into the same paragraph because the assignment is visible.
The roles can stay straightforward:
@evidence-agentfinds the source material and compares what the papers actually say.@methods-agentlooks for missing details that would make the work hard to reproduce.@tables-agentturns approved results into a LaTeX table, then makes sure it actually fits and compiles.@skeptical-reviewerlooks for a claim that has become too broad, a denominator that disappeared, or a citation that does not quite carry the sentence.@copy-editorcomes last, after the scientific work is settled.
You do not need all of these on every paper. But separating the jobs is much better than treating an agent as an all-purpose coauthor. A good prose editor should not be deciding whether a subgroup result is credible. A literature agent should not be rewriting the conclusion.
Paperclip Comes Before the Sentence
The most dangerous moment in AI writing is usually when the model has written a good sentence and then goes looking for a citation to attach to it. That is backwards.
For literature-heavy work, I would put Paperclip upstream of Sundial. Paperclip gives an agent structured access to full-text papers, trial records, regulatory documents, and other research material through a CLI, SDK, and MCP. In practice, that means the evidence agent can search the actual source set, inspect a methods section or supplement, and extract the same field across a few papers before it proposes anything for the manuscript.
Say we are working on a review section about treatment toxicity. The evidence agent can find the primary studies, pull the cohort, treatment, follow-up, endpoint definition, and event rate into a short comparison note, then suggest a very limited edit to the relevant paragraph. At that point, Sundial is where we read the suggestion, decide if it belongs, and compile the paper. Paperclip is where we make sure the sentence starts from something real.
This is not just a technical preference. It changes the conversation with the agent. Instead of asking, “Can you write this section?”, I can ask, “What do these six papers actually show, where do they disagree, and what can we say without stretching the evidence?” Then we can decide whether the answer deserves a sentence in the paper.
Review the Diff, Then Look at the PDF
Sundial’s reviewable changes are the feature I care about most. The agent does not need the authority to silently rewrite the paper. It needs the ability to do a useful pass that a human can inspect quickly.
I would still review the project in two different ways.
First, I would review the diff. Did the agent stay in the assigned file? Did it quietly alter a conclusion or replace a reference? Does every new factual statement have a source note behind it? That is how you keep the work controlled.
Then I would look at the compiled PDF. A diff cannot tell you that the table now runs off the page, a figure label has vanished, or a paragraph reads strangely after its neighbors changed. A clean build does not prove a scientific claim either, but it tells you that the paper you are discussing is the paper the reader will see.
Start with the project
Put the LaTeX source, bibliography, figures, and a short notes file in the shared workspace. Pick the authoritative version before anyone starts asking agents to edit.
Assign one contained job
Name the file or paragraphs, the evidence boundary, the required output, and the places the agent is not allowed to touch. Ask for a suggested diff and source note.
Retrieve before drafting
Use Paperclip or another source-aware tool to inspect the primary material first. The source should lead to the sentence, not the other way around.
Accept selectively and compile
Keep only the changes that survive author review, then build the PDF. Resolved references, readable tables, and a coherent manuscript are part of finishing the pass.
Do one human integration pass
Someone has to read the whole argument from start to finish. Scoped agents make that pass cleaner. They do not replace it.
The Boundary Is Not Optional in Medical Work
This setup is especially useful in academic medicine, where the work is collaborative and the stakes are high. A draft may include unpublished analysis, collaborator data, or patient-derived information. A convenient editor does not make any of that appropriate to upload somewhere.
Use institution-approved tools for protected information. De-identify when needed. Be clear about which source folders an agent can access. Keep journal policy and AI disclosure requirements in view. And do not confuse a signed suggestion with a responsible scientific author.
The people on the author line still own the study design, analysis, interpretation, and final claims. I would let agents retrieve, organize, compare, render, and propose. That is already a lot of help.
Why I Think This Is Worth Trying
I do not think the future of research writing is one agent sitting beside a blank page, ready to generate a manuscript on command. That produces too much language and not enough thought.
The better version looks more like a decent research team. Everyone works from the same project. People own parts of the manuscript. Agents get specific assignments. Sources are close enough to inspect. Changes are visible. The paper is rendered often enough that small technical problems stay small. Then a lead author pulls the whole thing together.
Sundial is appealing because it gives that workflow a home in the manuscript itself. Pair it with Paperclip for source-grounded research, and AI starts to feel less like a text generator and more like a useful colleague who can take a well-defined pass, show their work, and wait for you to decide.
