New in Beaver 0.24

batch jobs

Sort or tag a thousand items in a single request. Annotate, summarize or extract findings across hundreds of papers. Beaver keeps a ledger of every item, so the work is finished — and reported — item by item.

Filed items
680 items · not in any collection
Completed

File each item into the existing subject collection matching what the paper is about.

680 of 680 done100%
Where items went
Cancer & Immunology125
Neuroscience & Brain Research119
Cell & Developmental Biology109
Microbiology & Virology101
Genomics & Genome Engineering87
2 more collections · 607 filed in total
73 items examined and left unfiled — no existing collection was a genuine fit.

Work at the scale of your whole library

One request covers the entire population — every unfiled item, every untagged item, a whole collection, every record missing a DOI. Beaver works through it item by item and tells you exactly where each one ended up.

680
unfiled items sorted in a single request
across seven existing collections, none created
100%
of the population accounted for
every item filed, or recorded as needing no change
1 request
from asking to a finished job
Beaver keeps working until every item has a decision

How a batch job works

A batch job is a ledger of decisions, not a to-do list. Every item ends up either changed or recorded as needing no change — and Beaver cannot tick an item off by claiming it.

1

The whole population, resolved up front

Beaver names the job in one line — one operation, one set of items — and resolves the set straight from your library: everything unfiled, everything untagged, one collection, every record missing a DOI. The whole set is resolved in one step rather than a page at a time, so you see the exact count before anything runs.

2

You approve it before it starts

A batch job is shown to you as a card with the count, the goal and an estimated credit cost. Anything destructive — tags it will remove, values it will overwrite — is stated separately from the goal, so what you stand to lose never hides inside what Beaver intends to do.

3

Progress is measured, not reported

Each item is credited from the tool call that actually touched it. There is no “mark as done” step, because being done is derived from the ledger rather than claimed. A failed call credits nothing at all.

4

You watch the result take shape

A running tally shows where items are landing — 125 into Cancer & Immunology, 119 into Neuroscience — built from the calls themselves. Beaver is warned when one destination starts to swallow the batch, while there is still work left in which to change course.

5

It keeps going, and it can be resumed

Long jobs get a budget to match, and Beaver hands itself a fixed slice of items at a time, so a job of hundreds of items does not wander. If you stop it, the job is paused rather than lost — say “continue” and it picks up the same job instead of starting a new one.

6

You get a receipt, not a summary

When the last item has a decision, the job closes with a durable record: how many items went where, how many needed no change and why, what could not be read. The write-up you read is composed from those numbers, and the record stays in the conversation to answer follow-up questions later.

How much fits in one job

Limits are per job, and a job is one operation. A multi-stage request — tag, then sort — is one job per stage.

Sort into collections1,000 items
Tag1,000 items
Fix metadata1,000 items
Read & extract180 documents
Annotate100 documents
Write notes100 documents

Benchmarks

Beaver is evaluated against a frozen 866-item Zotero library on the jobs researchers actually run. Each case below was run on the same model with the same budget in both arms — the only difference is whether batch jobs were available.

Coverage — the share of requested items that actually got the work

Counted from the library afterwards, not from what Beaver said it did. Items that received only a blanket label do not count.

Beaver 0.23 — without batch jobsBeaver 0.24 — with batch jobs
Sort every unfiled item into the collection it belongs in
680 items · reuse existing collections, create none
Give every untagged item consistent subject tags
616 items · one shared vocabulary
Highlight the key findings in every paper in a collection
11 papers · small enough to finish either way
0%25%50%75%100%
JobMeasure0.230.24
Sort into collections
680 unfiled items
Items filed — best of three runs332617
Items filed — worst of three runs0600
Share landing in the largest collection16–27%20–21%
Tag
616 untagged items
Items given a tag that tells them apart0615–616
Distinct subject tags used0–19–91
Annotate
11 papers in a collection
Papers annotated1111
Highlights created28–2936–44

What the comparison run does instead

Without batch jobs the same model finds the population correctly and then settles it differently each time: one run reported the size of the job and asked how to proceed, one filed 206 items and stopped, and one gave all 616 untagged items the single tag “subject:biology”. All three are defensible answers to an open-ended request — none of them is the finished library you asked for.

How this was measured

Each case runs against a frozen 866-item Zotero library, reset to the same state before every run, on Beaver’s default model with the same request budget in both arms — three runs per arm, mean shown. Coverage is read back from the library afterwards, never from what the assistant said it did, and an item only counts when it received a tag or collection that actually tells it apart from the rest: a run that gives everything the same label scores zero, however many items it touched.

What to ask for

You never ask for a batch job. You ask for the work, and Beaver opens one when the job is big enough to need it.

Sort everything that isn't in a collection into my existing collections by subject. Don't create new ones.
Add consistent topic tags to every untagged item in my library.
Go through my Inbox collection and split it into “read” and “skip” based on whether it's about urban inequality.
Add a note to every paper in my Microbiology collection summarizing what it found.
Highlight the key findings in each paper in this collection — two or three per paper, from the results, not the abstract.
A batch of records imported badly. Find every item missing a DOI or journal and fix what you can look up.

Getting the most out of it

Name the population, and be honest about its size

“All six or seven hundred unfiled items” works better than “a few of my unfiled items”. Beaver resolves the set from your library either way, but understating the scale makes it stop early — that was measurable in our evaluations.

Say what a good outcome looks like for one item

The goal is judged one paper at a time. “File it into the existing subject collection that matches what the paper is about” gives Beaver something it can hold a single item against; “tidy up my library” does not.

State the constraints you actually care about

“Don't create new collections”, “reuse my existing tags”, “leave anything you can't verify and tell me”. These become part of the job and are checked against every item, rather than being forgotten twenty items in.

One kind of work at a time

Tag, then sort, then write notes — three jobs, run one after the other. A single request mixing all three gives every item a blurrier target and a longer run.

Check the estimate before approving

The approval card carries an estimated credit cost, and approving a job raises the ceiling for it so it isn't interrupted halfway through. Big jobs genuinely cost more than one wrong call — that's what doing the work per item means.

If it pauses, just say continue

A job you interrupted is paused, not lost. “Continue” resumes the same ledger with the same denominator — asking again from scratch starts a second job over the same items.

Read the closing report — it is the receipt

The final counts come from the ledger: how many items went where, how many were examined and left alone, what failed. If something is missing, that number is where it shows up.

Every change still needs your approval

A batch job doesn't change how Beaver touches your library. Edits are shown before they are applied and can be undone, and anything a job will remove or overwrite is declared on the approval card before it starts.