Language models can help researchers search literature, synthesize evidence, and work through complex questions. But scientific work places particular demands: answers must stay grounded in evidence, models must preserve what evidence actually supports rather than broadening conclusions, and researchers need to verify outputs.
Allen AI observed this in how scientists use Asta, its agentic platform for scientific work. Instead of simple keyword searches, users bring substantial context and constraints—for example, asking Asta to compare approaches across literature while accounting for particular methods, populations, or settings. Many treat generated reports as working research artifacts rather than one-off answers.
To help scientists generate cited reports faster with a model they could download and run themselves, Allen AI tested whether a small, open model trained for scientific report generation could match the quality of proprietary models while reducing generation time and serving costs.
The result is AstaBrief 8B, which turns a research question and retrieved literature excerpts into a cited report. AstaBrief is available in Asta's Generate a Report feature as Fast mode alongside Claude-powered Thinking mode. Allen AI is also open-sourcing the model and training data.
Developing AstaBrief required tens of thousands of real research queries, citation-focused filtering, preference data, and a redesigned report-generation pipeline that writes full reports in one pass rather than section-by-section. The result is nearly an order-of-magnitude reduction in report generation time: Fast mode averages 51.1 seconds per report compared with 178.5 seconds for Thinking mode—about 3.5× faster.
Open weights allow institutions to run AstaBrief on their own infrastructure, which matters when research questions reveal sensitive or unpublished work. Allen AI is releasing an example workflow that researchers can adapt to create reports from their own PDFs, providing a starting point for local report generation.
Training Approach
Allen AI started from Qwen3-8B and focused most effort on post-training data, evaluation, and surrounding report-generation scaffolding. Rather than pursuing reinforcement learning—which can be unstable and expensive—the team used a simpler recipe built around supervised fine-tuning (SFT) and direct preference optimization (DPO). This approach is cheaper, more operationally manageable, easier to debug, and iterate on.
That made training data quality especially important. Rather than relying on complex optimization methods to compensate for noisy examples, the team spent much of the project figuring out how to generate, select, and filter examples that demonstrated the desired report-writing behavior.
For speed improvements, Allen AI trained AstaBrief to directly generate final reports in one pass given a user query and relevant retrieved snippets, bypassing the expensive snippet summarization and clustering stages the Claude-based Thinking mode uses and avoiding section-by-section writing. Notably, this was possible without sacrificing performance.
Data Collection and Preparation

The training pipeline began with real user queries submitted through the system described in Allen AI's paper “Synthesizing scientific literature with retrieval-augmented LMs” and ScholarQA, the framework underpinning Asta's Generate a Report feature.
Analysis of hundreds of thousands of Asta queries showed that expert researchers frequently supplied substantial context, multiple constraints, and relationships between concepts rather than short, keyword-style prompts. Recent Asta user studies surfaced other differences: some researchers are comfortable using models for ideation or experimentation, while others prefer narrower roles in synthesis, literature surveillance, or pattern-finding. Across those differences, participants want clearer source traceability, more visibility into model behavior, and greater control over context.
Allen AI filtered user logs for quality, relevance, and privacy, stripping out beta-tester and bot traffic, dropping queries too short to be meaningful, and using LLM-based filtering to catch non-English queries, non-scientific requests, and prompts containing personal information. This left a pool of 90K research-focused queries.
For SFT, Allen AI generated full-report target outputs from filtered queries using the multi-step ScholarQA pipeline. The pipeline retrieved relevant literature, organized material into sections, and used backing report-generating models to synthesize evidence into cited reports. The team drew on Claude 3.5 Sonnet, Claude 3.7 Sonnet, o3, o4-mini, and GPT-4.1. After quality filtering, this yielded 47K usable training examples.
DPO required different training data—pairs of reports with one preferred over the other. Allen AI built pairs from a separate subset of queries not used during SFT data generation. One report per query came from the existing ScholarQA pipeline, backed by Claude 3.5 Sonnet or 3.7 Sonnet. The competing report came from o3, o4-mini, DeepSeek-V3, or DeepSeek-R1.
Two judge models—GPT-4.1 and DeepSeek-R1—compared each pair and picked a winner. Allen AI ensured LLM judges aligned with human preferences (95% agreement) and only kept pairs where both judges agreed, yielding a cleaner preference set with less noise. After quality filtering, the final DPO dataset contained about 6K examples.
Using multiple generators and requiring agreement between two judges provided a relatively simple way to construct preference data without treating any single model's output or judgment as ground truth.
Evaluation and Metrics
Allen AI's main evaluation target was SQABench-CS2, a set of 200 user-written computer science research questions. The team tracked four metrics:
- Rubric score: measures how much necessary content the report covers
- Answer precision: measures whether each paragraph is relevant to the question
- Citation precision: measures whether each citation supports its attached claim
- Citation recall: measures whether claims are fully supported by provided citations
For the final model, Allen AI ran secondary evaluations including DeepScholarBench, a 63-query benchmark for long-form research synthesis from recent ArXiv papers, and pairwise evaluations against reports from the Claude-powered pipeline—both LLM-judged comparisons on SQABench-CS2 and a small human study.
Scientific Faithfulness
Citation support is only part of scientific faithfulness. A model can cite the right study and still make stronger claims than the study supports. This can happen subtly: turning a finding about a particular sample into a generic population claim, shifting past-tense results into present-tense statements that sound more universally true, or converting descriptive findings into recommendations for clinicians, policymakers, or researchers.
These generalizations are especially important for scientific report generation because each step can broaden the apparent scope of evidence without introducing an obviously false statement. A cited sentence may be technically related to its source while still overstating what researchers established. Allen AI's development metrics focused on relevance, coverage, and citation grounding; a richer evaluation should also test whether scientific report writers preserve the scope and strength of source claims.
Initial SFT runs improved content quality but still lagged behind the Claude-powered pipeline on answer precision and citation quality. The model improved at writing reports but wasn't grounded in evidence as consistently as needed for scientific synthesis.
Most Read
