Offshoot: Literature Orchestrator

How Offshoot is organized: a scheduled GitHub Actions run starts the coordinator, which uses stored passwords and a memory database, and hands jobs to three AI helpers: a collector, an idea finder and judge, and a writer. Each helper reports back to the coordinator; they never talk to each other. The collector gets new papers from arXiv, and the writer creates project drafts in Google Docs. GitHub Actions Secrets Memory arXiv Google Docs scheduled run database new papers project drafts Coordinator Collector Idea finder + judge Writer tracks progress
The coordinator is the manager. It hands each job to a specialist AI helper and collects the result; the helpers never talk to each other directly. The Collector gathers new papers, the Idea finder + judge picks out ideas and decides which fit, and the Writer turns them into drafts. Solid lines carry papers and results; dashed lines carry passwords and saved records.

Building an AI system that reads new research papers and turns their unanswered questions into ideas for my next research project.

Built as Offshoot. Status: in development.

Problem and motivation

Research where biology meets AI moves fast in two fields at once, and the best project ideas rarely appear in a paper's summary. They're buried in a line like “further studies are needed” near the end, or in a list of open questions. Most tools that track new research stop at summarizing what a paper found; Offshoot goes after what a paper left unfinished and turns it into a project idea I can act on.

Approach

Offshoot works like a small team. A coordinator, the manager, hands each job to a specialist AI helper and collects the results. Each helper gets exactly the information it needs, several can work at the same time, and the coordinator keeps track of what's been done and redoes anything that goes wrong.

Automatic start

It runs on a schedule by itself on GitHub's servers, so it doesn't depend on my laptop being on.

Collecting papers

It pulls new papers from arXiv, a free online archive of research papers, in categories where biology overlaps with AI. It always looks back over the same stretch of recent submissions, so if one run is missed, the next one catches up.

Finding and judging ideas

An AI helper pulls out each paper's stated limitations, suggestions for future work, and methods that could be reused elsewhere. Its answers have to follow a set format, and they're checked and redone if they come out wrong. Each idea then gets a simple yes or no on whether it fits my research, based on examples I labeled by hand and one written set of criteria, so the standard stays consistent.

Writing it up

Accepted ideas become project drafts in Google Docs, with a goal, the gap in the research it's based on, steps, size and difficulty labels, and a link to the original paper. Rejected papers are saved for later review, and ideas the system is unsure about are flagged in the draft.

Memory and security

A small database remembers every paper it has already seen, so nothing gets sent twice. The code will be public on GitHub, so passwords and keys are kept out of it, and its access to Google is limited to only the files it works with.

Outcome and results

The overall design is settled: which papers to collect, the yes-or-no judging, sending drafts only to Google Docs, and how the code is organized. I built a simple working version first, with one helper collecting papers, one pulling out ideas, and drafts written up directly, before adding extra features.

So far it has surfaced 50 papers and produced 5 possible project drafts. One of those became a project I've started: using continuous-time recurrent neural networks (CTRNNs), a type of AI model inspired by how neurons behave over time, to study immune network theory, the idea that parts of the immune system regulate one another as a network.

Skills practiced

  • Designing a team of AI helpers that work together under a coordinator
  • Getting AI to give answers in a fixed, checkable format, and retrying when it doesn't
  • Building an automatic process that avoids duplicates and recovers from missed runs
  • Keeping passwords and keys out of public code

Challenges and what I'd do next

Keeping the scope tight: an earlier design also added tasks to my to-do list and events to my calendar automatically. I cut that, because deciding to commit to a project should be my call during my weekly review, not the system's.

Next, on top of the working version: processing papers in batches, handling uncertain results more carefully, and splitting “finding ideas” and “judging fit” into separate steps. With that split, a retry fixes the actual problem: a badly read paper gets re-read, while a paper that simply isn't a good fit doesn't get re-read for nothing.

Tools

  • Claude Agent SDK (AI helpers)
  • arXiv (research papers)
  • SQLite (database)
  • Google Docs
  • GitHub Actions (scheduling)
  • Python

Thank you to arXiv for use of its open access interoperability.