Bart: A Vintage LLM

Spent so far$0.00

#Introduction

"The farther backward you can look, the farther forward you are likely to see." — Winston Churchill

We can't fully grasp modern questions like the rise of artificial intelligence, economic disruption, or automation while standing in the middle of them. Sometimes our best path forward is the one that takes us backward. In the past year, many have begun to explore this idea by applying it to large language models (LLM's). What can we learn about today's modern technology when we fuse it with information hundreds of years old? Recent models such as Talkie,Radford, Levine and Duvenaud. talkie — a 13B model trained on 260B tokens of pre-1931 English. talkie-lm/talkie-1930-13b-base and GPT-1900,Hla, M. GPT-1900 — a 3.3B model trained from scratch on pre-1900 text, to see whether it could arrive at quantum mechanics and relativity. michaelhla/gpt1900 have demonstrated that stripping away modern hindsight transforms an artificial mind from a simple archive into a powerful diagnostic lens. But while these models offer a proof of concept, true historical reasoning—like devising the transistor 20 years ahead of schedule—is still unreached. Material results remain the holy grail of historical AI. The next step, then, is to push historical LLM's into the future: state-of-the-art architectures, cleaner, more refined corpora, period-appropriate evaluations, and vintage-adapted post-training techniques. In this article, we introduce Bartholomew III, or Bart for short, our attempt to move the needle of historical language modeling.

#Results

Bart is a depth-32, 2.82B-parameter decoder-only transformer trained on 20.1B tokens of pre-1931 English. Every score below is Vintage CORE — CORE with its post-1930 knowledge removed, described under Evaluation — on the twenty tasks common to all three bundles.

Vintage CORE profile
Mean centered score across the 20 common tasks, grouped into seven domains · Restyled bundle
0.15 0.30 0.45 0.60 Language modeling 1 task Reading comprehension 3 tasks Commonsense 5 tasks Reference knowledge 3 tasks Academic QA 4 tasks Coreference 2 tasks Symbolic reasoning 2 tasks 0.30 0.60 Language 1 task Reading 3 tasks Commonsense 5 tasks Reference 3 tasks Academic QA 4 tasks Coreference 2 tasks Symbolic 2 tasks

Pick a model below to bring it forward. Bart's shape is a vintage model's shape: it holds its own at language modeling and coreference, and falls away where world knowledge is needed. Talkie, at 13B and 260B tokens, is the outer envelope.

Data
ModelLanguage modelingReading comprehensionCommonsenseReference knowledgeAcademic QACoreferenceSymbolic reasoning
Bart0.4140.1380.2130.2080.1200.1980.064
Bart d240.3530.0810.1710.1700.0830.1640.057
GPT-1900 d340.401-0.0370.1580.2030.0880.2880.080
Talkie 1930 13B0.6110.4100.3400.3920.2200.4650.153
Modern d240.4260.1390.3110.3100.2610.2620.086
Every model across the seven CORE domains, on the restyled bundle. Pick a model in the legend to bring it forward — Bart starts in front.
Data — every benchmark, every model

Every task in the common-20 intersection, grouped by domain. Scores are CORE's centered accuracy — (accuracy − random) / (1 − random) — so 0 is chance and a negative number means the model did worse than guessing.

Every task in the common-20 intersection, grouped by domain. Scores are CORE's centered accuracy — (accuracy − random) / (1 − random) — so 0 is chance and a negative number means the model did worse than guessing.

Original
BenchmarkBartBart d24GPT-1900 d34Talkie 1930 13BModern d24
Language modeling
LAMBADA0.4140.3530.3970.6070.427
Reading comprehension
SQuAD0.2030.1140.1270.5530.354
CoQA0.2660.1660.2180.4310.255
BoolQ-0.060-0.061-0.3990.281-0.092
Commonsense
HellaSwag zero-shot0.1660.0980.1240.3560.356
HellaSwag0.1710.1050.1240.3880.357
CommonsenseQA0.0860.1340.0530.0750.147
COPA0.3600.2200.1800.4600.340
PIQA0.2080.1610.1620.3660.428
Reference knowledge
Jeopardy0.0330.0040.0330.3330.193
BIG-bench Wikidata QA0.3170.2520.2870.4810.511
BIG-bench language ID0.1770.1830.1830.1790.174
Academic QA
ARC Easy0.2780.2230.2170.4830.570
ARC Challenge0.025-0.0140.0010.1820.181
OpenBookQA0.0450.0110.0400.1200.200
LSAT analytical reasoning0.0870.0760.0820.0650.130
Coreference
Winograd0.2970.3410.3920.6410.421
WinoGrande0.066-0.0120.0800.2800.133
Symbolic reasoning
BIG-bench operators0.1290.1140.1290.2430.171
BIG-bench repeat/copy logic0.0000.0000.0000.0620.000
Filtered
BenchmarkBartBart d24GPT-1900 d34Talkie 1930 13BModern d24
Language modeling
LAMBADA0.4150.3520.4010.6110.425
Reading comprehension
SQuAD0.2130.1190.1390.5420.348
CoQA0.2820.1780.2360.4290.238
BoolQ-0.078-0.012-0.4670.262-0.177
Commonsense
HellaSwag zero-shot0.1880.1150.1440.3840.354
HellaSwag0.1960.1290.1480.4160.357
CommonsenseQA0.0610.1260.0500.0650.155
COPA0.3800.2400.1600.5000.340
PIQA0.2770.2300.1930.4270.422
Reference knowledge
Jeopardy0.0420.0050.0480.3980.190
BIG-bench Wikidata QA0.4080.3310.3790.6160.562
BIG-bench language ID0.1680.1720.1770.1750.177
Academic QA
ARC Easy0.3260.2440.2470.5430.571
ARC Challenge0.042-0.0010.0100.2250.184
OpenBookQA0.0690.0350.0610.1570.229
LSAT analytical reasoning0.0110.0110.011-0.0050.103
Coreference
Winograd0.2820.3330.3850.6260.436
WinoGrande0.091-0.0060.0970.3070.143
Symbolic reasoning
BIG-bench operators0.1290.1140.1290.2430.171
BIG-bench repeat/copy logic0.0000.0000.0000.0620.000
Restyled
BenchmarkBartBart d24GPT-1900 d34Talkie 1930 13BModern d24
Language modeling
LAMBADA0.4140.3530.4010.6110.426
Reading comprehension
SQuAD0.2050.1100.1370.5320.344
CoQA0.2730.1730.2330.4160.234
BoolQ-0.065-0.038-0.4800.281-0.161
Commonsense
HellaSwag zero-shot0.1880.1180.1500.3730.344
HellaSwag0.1960.1300.1510.3970.336
CommonsenseQA0.0680.1380.0470.0670.152
COPA0.3600.2400.2400.4400.320
PIQA0.2520.2280.2010.4240.401
Reference knowledge
Jeopardy0.0460.0070.0510.3850.190
BIG-bench Wikidata QA0.4080.3310.3790.6160.562
BIG-bench language ID0.1680.1720.1770.1750.177
Academic QA
ARC Easy0.3310.2460.2500.5200.538
ARC Challenge0.066-0.0010.0180.2370.180
OpenBookQA0.0720.0770.0720.1310.221
LSAT analytical reasoning0.0110.0110.011-0.0050.103
Coreference
Winograd0.3110.2820.4580.6410.407
WinoGrande0.0840.0470.1180.2880.118
Symbolic reasoning
BIG-bench operators0.1290.1140.1290.2430.171
BIG-bench repeat/copy logic0.0000.0000.0310.0620.000

#Leakage

Historical events become more surprising after model cutoffs

HISTORY-EVENT independent reconstruction · decade macro mean · bands show ±1 SE

GPT-1900 cutoff  1900 Bart / Talkie cutoff  1930
HISTORY-EVENT bits per byte by event decade Seven model series from 1700 through 2020. Historical language models rise after their cutoff years, while the modern SmolLM3 proxy stays comparatively flat.
Lower BPB means an event is less surprising to the model. Hover to compare models at any decade.
Bits-per-byte on HISTORY-EVENT by event decade. Lower BPB means an event is less surprising to the model. Hover to compare models at any year.

The dataset is an independent reconstruction of the HISTORY-EVENT benchmark described in Pretraining Language Models on Historical Text.arXiv:2606.02991. This is not the authors' official dataset; our public reconstruction is at jbduran/history-event-reconstruction. We can see an increase after the knowledge cutoff starting around the 1900s for vintage models. In contrast, SmolLM3 remains broadly flat across decades, consistent with modern-event exposure.

#Pretraining

Arguably, the most important part of a model is its dataset. Curating a high quality dataset means obtaining raw pre-1930s text that's clean, correctly dated, and large enough to pretrain on. Combining these constraints leads to significant challenges.

Of the three, we felt that text-cleanliness was the limiting factor in dataset curation and therefore was a natural place to begin. Optical character recognition (OCR) quality is a score representing how accurately text is captured from real-world images and scans. This is important because pretty much all text used in a historical model will have been sourced via OCR. Here is an example of poor OCR:

A scanned newspaper column beside its OCR transcription. The transcription garbles names and numbers, breaks words across lines, and inserts stray punctuation.
A representative low-confidence passage. Scanner noise, broken ligatures, and page furniture survive transcription intact.

For us, OCR quality is measured as a per-document confidence score. Low-scoring documents carry systematic transcription errors and stray page furniture that, if left in, corrupt the corpus at the subword level: the tokenizer allocates vocabulary to garbage merges, and the model spends capacity modeling scanner noise instead of language. One possible way of mitigating OCR issues would be to re-OCR pre-1930's texts via a modern LLM-based OCR, but after considering this approach we found that it would cost too much. Assuming roughly 16–20 million pages for an 8–10B token dataset at around 500 tokens per page, re-OCRing even a small dataset would cost about $16,000 at Mistral OCR batch pricing.Mistral AI. Introducing Mistral OCR 3 — a vision-language OCR model with a batch API discount. mistral.ai/news/mistral-ocr-3

Alternatively, we found it best to source pre-existing datasets and transform the existing corpora. A few options stood out. Project Gutenberg is a high quality dataset but is also very small.gutenberg.org — 75,000+ hand-proofread public-domain books. The American Stories dataset had great size, but we were concerned with the OCR quality due to the difficulty of OCR-ing newspapers.Dell et al. American Stories: A Large-Scale Structured Text Dataset of Historical U.S. Newspapers, NeurIPS 2023. arXiv:2308.12477 · dell-research-harvard/AmericanStories We kept searching, and came across Institutional Books 1.0. This dataset is a result of the digitization and post-processing work of the Harvard Library, boasting 242B tokens and almost a million documents.Cargnelutti et al. Institutional Books 1.0: A 242B Token Dataset from Harvard Library's Collections, 2025. arXiv:2506.08300 · institutional/institutional-books-1.0

Working from the Institutional Books (IB) dataset, we applied an aggressive OCR score threshold of 0.90, drawing on both OCRoscope and the OCR score from Google Books metadata, and required that the two metrics agree within 0.10 on any given document. Furthermore, we filtered by date (any undated books are rejected), we filtered for English (this resulted in the most rejects), and we filtered text with tokenizability scores less than 95 — a score indicating how efficiently a tokenizer encodes the text. This trimmed the 242B tokens down to 27B. We then deduplicated by book barcodes and shuffled the dataset. The result was our first cleaned dataset — bart-dataset-v1.

Bart dataset by subject
Share of the retained documents, by Library of Congress class
28.6% 18.4% 8.6% 8.3% 6.0% 5.4% 160,263 documents Language and literature 28.62% Philosophy, psychology, religion 18.44% Law 8.61% Science 8.30% History of the Americas 5.95% Social sciences 5.37% Auxiliary sciences of history 3.37% Agriculture 3.30% Political science 3.10% Education 2.73% Technology 2.33% Geography, anthropology 2.14% Fine arts 1.99% Medicine 1.89% Other (7 subjects) 3.85% 28.6% 18.4% 8.6% 8.3% 160,263 documents Language and literature 28.62% Philosophy, religion 18.44% Law 8.61% Science 8.30% History of the Americas 5.95% Social sciences 5.37% Auxiliary history sci. 3.37% Agriculture 3.30% Political science 3.10% Education 2.73% Technology 2.33% Geography, anthropology 2.14% Fine arts 1.99% Medicine 1.89% Other (7 subjects) 3.85%

The seven classes below 1% are grouped; every one of them is listed in the data table. Science, technology and medicine together come to 12.5% — the shortfall the midtraining mixture was built to correct.

Data
SubjectDocumentsShare
Language and literature45,86328.62%
Philosophy, psychology, religion29,54818.44%
Law13,7988.61%
Science13,3088.30%
History of the Americas9,5435.95%
Social sciences8,6035.37%
Auxiliary sciences of history5,4043.37%
Agriculture5,2903.30%
Political science4,9663.10%
Education4,3812.73%
Technology3,7332.33%
Geography, anthropology3,4332.14%
Fine arts3,1921.99%
Medicine3,0341.89%
Music1,5730.98%
World history1,4810.92%
Naval science1,0000.62%
General works9140.57%
Military science8290.52%
Bibliography, library science3210.20%
Unknown490.03%
Literature, philosophy and religion take almost half the corpus between them. Science, technology and medicine — what we most wanted the model to know — total 12.5%, which is what later drove the midtraining mixture.

Another phase of filtering targeted boilerplate text and OCR corruption that snuck through IB's filters. This consisted of removing useless headers, footers and library stamps, alongside gibberish symbols and antiquated characters through heuristic regex filters. We also followed Michael Hla's log-prior filter, which estimates a text sequence's unconditional log-prior probability by summing token-level log probabilities and then trims improbable or out-of-distribution tokens.Documents whose log prior falls outside the 2.5th–97.5th percentile band are dropped, trimming both ends of the distribution. See michaelhla/gpt1900. This phase resulted in bart-dataset-v2. While we expected major improvements in our validation bpb (see Evaluation), our models actually saw a decline in performance because this phase removed most of the easy items for a model to replicate — that is, the boilerplate.

Thrown out

Copyright (c) 1966, by Macmillan Publishing Company, a division of Macmillan, Inc.

Library of Congress Catalog Card Number: 66-16140 Printed in the United States of America

ISBN O0-Oe-33ec00-4 Printing 30 29 28 Year 0

ISBN O-Oe2-33ece00-4

Kept

Two bodies of unequal weight (say a guinea and a feather) are placed at the same height under the exhausted receiver of an air-pump. When released, they are observed to reach the bottom of the vessel at the same instant of time, or, in other words, to fall in equal times. From this fact, it is inferred that a repetition of the experiment either with these two bodies or with any other bodies would be attended with the same result.

The same filter, both outcomes. On the left a copyright page that teaches the model nothing but its own format; on the right a passage of pre-1930s physics.

For our final phase, we decided to focus on the time period. Although we did filter previous stages by date, leaks and bad labeling would inevitably bring modern documents to contaminate our vintage dataset. We once again filtered through the boilerplate of documents, focusing on removing missed footers. However, the biggest part of this phase was the regex anachronism filter. We used a tiered banned words list to flag documents and remove them on too many "hits."

Tier 1 1 hit discards
  • smartphone
  • world war ii
  • dna
Tier 2 2 hits discard
  • email
  • video game
  • nuclear power
Tier 3 only counts beside a Tier 2 hit
  • computer
  • satellite
  • drone
Three example terms from each tier of the anachronism list. The real list is considerably longer.

Rather than trimming, we removed entire documents when conditions were met. A single Tier 1 hit was enough to discard a document, while Tier 2 words required two occurrences. Tier 3 words never triggered removal on their own — they only counted when accompanied by at least one Tier 2 hit. We used this strategy later in midtraining to keep all of our data as vintage as possible. With all of these strategies, we created bart-dataset-v3 — our best vintage dataset yet.

Token mass removed at each cleaning phase
0 5 10 15 20 25 30 29.69B bart-dataset -v1 −0.80B 2.7% OCR artifacts −1.14B 3.8% Prior low p<2.5 −1.18B 4.0% Prior high p>97.5 26.57B bart-dataset -v2 <0.01B Footer strip −0.87B 2.9% Anachronism filter 25.70B bart-dataset -v3 Tokens remaining (billions) 0B 10B 20B 30B bart-dataset-v1 29.69B OCR artifacts −0.80B Prior low −1.14B Prior high −1.18B bart-dataset-v2 26.57B Footer strip <0.01B Anachronism −0.87B bart-dataset-v3 25.70B Tokens remaining (billions)

Derived from exact character counts: 118.7G → 106.3G → 102.8G characters, at 4.0 characters per token. End-to-end retention is 86.57% of tokens and 91.12% of documents.

Data
StepTokensRemovedShare
bart-dataset-v129.69B
OCR artifacts−0.80B2.7%
Prior low p<2.5−1.14B3.8%
Prior high p>97.5−1.18B4.0%
bart-dataset-v226.57B
Footer strip<0.01B<0.01%
Anachronism filter−0.87B2.9%
bart-dataset-v325.70B
Token mass removed at each phase of the cleaning cascade. The two log-prior trims cost more than every other filter combined.

#Architecture

Every successful model builds on a proven foundation and modifies it for its needs. When Qwen trained their first model family (Bai et al.),Bai et al. Qwen Technical Report, 2023. arXiv:2309.16609 they started from Llama's architecture. Kimi K2 started from DeepSeek-V3's MoE architecture.DeepSeek-AI. DeepSeek-V3 Technical Report, 2024. arXiv:2412.19437 We used Karpathy's Nanochat as a base since it's highly optimized and easily modifiable.karpathy/nanochat — an end-to-end LLM pipeline in a single readable codebase. Nanochat is an end-to-end LLM pipeline (extending from tokenization all the way to inference) built around a single complexity dial — "depth". Depth drives model scaling and is what we use to research the model at different sizes. We abbreviated it to "d#", where the number corresponds to the chosen depth (eg. d12, or d24). Within the codebase is the actual GPT module: a decoder-only transformer with rotary embeddings, QK-norm, untied embedding/lm_head, ReLU² MLP, Group-Query Attention, and Flash Attention 3.Shah et al. FlashAttention-3: Fast and Accurate Attention with Asynchrony and Low-precision, 2024. arXiv:2407.08608

Model architecture
Bart — a decoder-only Transformer, after nanochat
32 layers·2,048 model width·16 × 128-dimensional heads·2.819B parameters
S · 1,024-token causal window L · full 4,096 context mid-stack backout
Token IDs 32,768-token vocabulary Embedding + RMSNorm 2,048 wide · learned smear 32 Transformer blocks 24 short-window · 8 full-context · 16 with value embeddings S S S L × 8 · 1,024-token window, then full 4,096 inside each pre-norm block Residual conditioning λr · x + λ0 · x0 Causal self-attention QKV · RoPE θ=100k · QK-norm FlashAttention-3 · optional VE MLP + residual 1× → 4× → 1× squared ReLU, bias-free Mid-stack backout subtract the cached layer-17 state RMSNorm → untied LM head softcap 15 · 32,768 logits Token IDs 32,768-token vocabulary Embedding + RMSNorm 2,048 wide · learned smear 32 Transformer blocks 24 short · 8 full · 16 VE S S S L × 8 inside each pre-norm block Residual conditioning λr · x + λ0 · x0 Causal self-attention RoPE θ=100k · QK-norm FlashAttention-3 MLP + residual 1×→4×→1× squared ReLU Mid-stack backout subtract the cached layer-17 state RMSNorm → untied LM head softcap 15 · 32,768 logits

The stack repeats a three-short-one-full attention pattern eight times. Value embeddings are learned on every other layer, and the backout subtracts the cached layer-17 state before the head.

The d32 production model, end to end. The block interior is drawn inside the stack it repeats, and the dashed line is the backout reaching back for the cached layer-17 state.

While we started with almost all of nanochat's hyperparameters, we wanted to verify that they were optimal for our vintage data. Through later experimentation, we tested countless different hyperparameters, most notably context size and the Muon optimizer's momentum.

#Evaluation

We knew where our data was coming from, and how it would be processed, but before we could train anything we needed to know how to gauge progress. Think about it: how exactly do you quantify a good vintage model?

Validation BPB is a way of seeing how good a model is at autocompleting a set-aside, unseen piece of the dataset. Validation BPB is a great metric for pretraining, but it is dataset-specific. Consequently, we're not able to compare against other vintage or modern LLMs.

So we reached for CORE, the benchmark suite nanochat reports for its GPT-2 speedruns.Li et al. DataComp-LM: In search of the next generation of training sets for language models, 2024. arXiv:2406.11794. CORE averages a centered score over 22 tasks: (accuracy − random) / (1 − random), so 0 is chance and 1 is perfect. CORE is dataset-independent, cheap, and comes with a known GPT-2 reference point. One problem: CORE assumes a modern world model. Its tasks ask about websites, post-1930 organizations, brands, later scientific discoveries, and contemporary sports. Is our model dumb, or is it just being asked modern-laced questions?

We remove two benchmarks that cannot be salvaged, filter post-1930s content, restore small tasks to size with reviewed replacements, and build a second bundle where the surviving prose is rewritten into the style of 1800–1930. The result is Vintage CORE, released as a versioned dataset alongside its evaluator. That gives three measurements:

  • Original — CORE as published.
  • Filtered — post-1930 knowledge removed. Original prose.
  • Restyled — same as filtered except vintage prose.
Vintage CORE model comparison
Every model on one scale · each row spans its three bundles
Original Filtered Restyled
0.10 0.15 0.20 0.25 0.30 0.35 0.40 Talkie 1930 13B 0.329–0.349 Modern d24 0.253–0.263 Bart 0.163–0.176 Bart d24 0.123–0.137 GPT-1900 d34 0.121–0.139 Common 20-task centered CORE score — higher is better 0.10 0.20 0.30 0.40 Talkie 1930 13B Modern d24 Bart Bart d24 GPT-1900 d34 Centered CORE — higher is better

Filtering lifts every vintage model and leaves the modern baseline flat or slightly worse, which is the result the benchmark was built to be able to show. Original is recomputed on the same 20 tasks as Filtered and Restyled.

Data
ModelOriginalFilteredRestyled
Talkie 1930 13B0.3290.3490.342
Modern d240.2630.2620.253
Bart0.1630.1750.176
Bart d240.1230.1360.137
GPT-1900 d340.1210.1270.139
All five models on one scale. Filtering lifts every vintage model and leaves the modern baseline flat or slightly worse — which is the result we wanted the benchmark to be able to show.
Data — headline scores
ModelCorpusDepthTraining tokensOriginalFilteredRestyled
BartPre-19313220.1B0.1630.1750.176
Bart d24Pre-1931248.76B0.1230.1360.137
GPT-1900 d34Pre-19003422B0.1210.1270.139
Talkie 1930 13BPre-1931260B0.3290.3490.342
Modern d24Modern248.76B0.2630.2620.253

The same broad strategy appears in TypewriterLM's HellaSwag-1800: filter the temporally unfair items, then rewrite the survivors.Pretraining Language Models on Historical Text, 2026. arXiv:2606.02991 · typewriter.chat We apply it across a full 20-task suite rather than a single benchmark.

#Ablations

In our hopes to push the historical modeling space forward, we're including a short space here to document our exploration of the many mechanics and quirks of vintage language models. We explored model behavior through ablations (aka experiments). The goal of ablations was to run experiments at a small scale and get results we could confidently extrapolate to our Bart model. We tracked most of our ablations on GitHub via Issues.Every ablation has an issue with its config, its W&B run, and its result: zachnorton14/bart/issues. Below we list some of the more interesting ablations that guided our model's trajectory. Every run behind them — 39 in all, with their tokenizers, checkpoints and evaluations — is archived in bart-experiments. We hope this all helps you in your own development.

Briefly, we want to also emphasize the importance of two key development tools we used: Weights and Biases, and model configs. Weights and Biases is a diagnostic service that tracks all things regarding model training.wandb.ai — learning rate, FLOPs, even GPU wattage. Model configurations are simply JSON files that define model hyperparameters for each run. While often unspoken of in ML engineering, these two things are essential for scalable and easy experimentation.

All ablations · validation BPB
Every run in the W&B project, 29 in all · axis clipped at 2.0 — every run starts near 3.41 at step 0 · lower is better
think (raw, d12) thinkcleaned (d12) clean1930s (d12) 1930s (d12) pre1900 (d12) ClimbMix comparison d24 runs Final d32
0.8 1.0 1.2 1.4 1.6 1.8 2.0 0 1k 2k 3k 4k 5k 6k 7k 8k 9k Training step Validation bits-per-byte 0.8 1.0 1.2 1.4 1.6 1.8 2.0 0 3k 6k 9k Training step

The d12 arms are coloured by the corpus their run name leads with; hovering a curve thickens it and names the exact run. The ClimbMix arm is the lowest depth-12 curve — the modern-data comparison from EXP005. The two long runs are the d24 and the final d32, which appears in W&B as two runs split at step 5500 by the midtrain-shard fix.

Data
RunGroupLast stepFinal val/bpb
Think.Unbounded-d32-v2mix-contfinal d329,6000.771
climbmix-d12-1epoch-25shardsClimbMix comparison (d12)2,5200.879
clean1930s-d24-r12-ctx4096-sssl-fulltok-mix-ogd24 run8,3520.886
Think.Unbounded-d32final d325,5000.900
clean1930s-d24-r12-ctx4096-sssl-fulltok-v1d24 run8,3520.920
think-d12-r11.25-ctx8192think · raw corpus (d12)2,3621.040
think-d12-1ep-65sh-r30think · raw corpus (d12)6,3001.060
think-d12-r20-2epochthink · raw corpus (d12)4,9951.062
think-d12-r20-alt44think · raw corpus (d12)4,2001.074
think-d12-1ep-44sh-r20-wd42think · raw corpus (d12)4,2001.078
think-d12-r11.25-ctx4096think · raw corpus (d12)2,3621.080
1930s-d12-r12-4096ctx-run21930s (d12)2,5201.088
clean1930s-d12-r11.25-ctx4096-sssl-fulltok-artransfer-muon090-s42-v1clean1930s (d12)2,3621.089
1930s-d12-r11.25-4096ctx-run11930s (d12)2,3621.089
clean1930s-d12-r11.25-ctx4096-full-fulltok-ablation-v1clean1930s (d12)2,3621.090
clean1930s-d12-r12-ctx4096clean1930s (d12)2,5201.090
clean1930s-d12-r11.25-ctx4096-sssl-fulltok-artransfer-muon083-s42-v1clean1930s (d12)2,3621.091
clean1930s-d12-r11.25-ctx4096-sssl-fulltok-artransfer-baseline-s42-v1clean1930s (d12)2,3621.092
clean1930s-d12-r11.25-ctx4096-sssl-fulltok-ablation-v1clean1930s (d12)2,3621.092
thinkcleaned-d12-r11.25-ctx4096thinkcleaned (d12)2,3621.093
think-d12-r11.25-run3think · raw corpus (d12)2,3621.102
think-d12-r11.25-run2think · raw corpus (d12)2,3621.102
think-d12-1ep-25sh-r11-wd42think · raw corpus (d12)2,3621.104
clean1930s-d12-r11.25-randtokclean1930s (d12)2,3621.113
thinkcleaned-d12-1ep-sh26-r11thinkcleaned (d12)2,3621.115
pre1900-d12-1ep-4sh-r20pre1900 (d12)4,2001.116
clean1930s-d12-r11.25-muoneq-muonplusclean1930s (d12)2,3621.116
clean1930s-d12-r11.25clean1930s (d12)2,3621.118
think-d12-r11.25-ctx4096-ssslthink · raw corpus (d12)2501.535
Validation bits-per-byte for every run in the project. The lowest depth-12 curve is the ClimbMix comparison run; the two runs that continue to the right are the d24 and the final d32.

#EXP007 · Variance testing

When testing on a small stand-in model, we found that CORE scores varied significantly, even for identical models, which, unfortunately, threw off our judgments of multiple previous experiments. A valuable reminder to not take anything for granted. While there is significant variation for CORE scores at the d12 scale, the amount of variation shown in EXP005 was unprecedented.

Run-to-run variance at 2k and 4k context
Bart d12 · three independent runs per context length · same eval suite
2k context (2048 tokens) 4k context (4096 tokens)
CORE score higher is better 0.060 0.065 0.070 0.075 0.080 0.085 0.0715 sd 0.0067 2k context 2048 tokens · n = 3 0.0742 sd 0.0072 4k context 4096 tokens · n = 3 Validation bits-per-byte lower is better 1.030 1.035 1.040 1.045 1.050 1.055 1.0513 sd 0.0008 2k context 2048 tokens · n = 3 1.0345 sd 0.0030 4k context 4096 tokens · n = 3 CORE score higher is better 0.060 0.065 0.070 0.075 0.080 0.085 0.0715 sd 0.0067 2k context 0.0742 sd 0.0072 4k context Validation bits-per-byte lower is better 1.030 1.035 1.040 1.045 1.050 1.055 1.0513 sd 0.0008 2k context 1.0345 sd 0.0030 4k context

Each dot is one training run. The capsule spans the group min–max, the horizontal tick is the group mean, sd is the sample standard deviation.

Data
RunContextCOREVal BPB
think-d12-r11.25-run12k context0.07921.0520
think-d12-r11.25-run22k context0.06691.0514
think-d12-r11.25-run32k context0.06831.0505
clean1930s-d12-r12-ctx40964k context0.08081.0324
1930s-d12-r11.25-4096ctx-run14k context0.06651.0380
1930s-d12-r12-4096ctx-run24k context0.07521.0332
Three identical runs per context length, seeds aside. The CORE spread is wide enough to swallow most of the differences we had been reading as results; validation BPB is far steadier.

#EXP005 · Modern and vintage mixture

A simple test to explore just how much our performance depends on dataset quality. Even a small inclusion of the heavily-optimized ClimbMix dataset late into the training schema — the last 20% of train shards — saw an enormous CORE boost: 0.0673 → 0.1003. This confirmed a suspicion we had carried for some time: our model's dominant bottleneck was the absence of modern data, not the quality of the vintage data we had curated.

Training input vs. Vintage CORE
Training tokens are a proxy for input volume · Original bundle, common 20-task centered CORE
Same d24 input: Modern 0.263 vs. Bart d24 0.123 at 8.76B tokens. Token scale: Talkie uses 6.9× Modern d32's input for 0.008 more CORE.
Token count alone is not FLOPs: model size, architecture, context length, hardware, and corpus composition differ. GPT-1900 (~22B) and Talkie (260B) are model-card figures; Modern d32 scores are reported to four decimals.
Training tokens against Vintage CORE. At identical input — 8.76B tokens, depth 24 — the modern corpus scores 0.263 and the vintage one 0.123. The gap is the corpus, not the recipe.

#EXP010 · Longer context length

Since the average pre-1930's document length is hundreds of times larger than that of modern pretrain corpora, increasing context length early in the training pipeline could benefit from these larger text sequences.This experiment had a flaw worth noting: it used full attention rather than sliding-window attention. Surprisingly, performance gains were merely cosmetic, and early tokens in the attention pattern showed minimal improvement in learned representation. The benefits still outweighed the costs, and we kept a 4096 context length.

Context length: what each doubling costs and buys
Cost versus performance 1.00 1.02 1.04 1.06 1.08 1.10 1.12 1.14 100 110 120 130 140 150 160 ctx2048 1.11 bpb ctx4096 1.08 bpb ctx8192 1.04 bpb Runtime (minutes) — cost val/bpb Marginal efficiency of each doubling 0.0 0.5 1.0 1.5 2.0 2.5 3.0 2.13 +14.1 min −0.03 bpb 2k → 4k 1.32 +30.2 min −0.04 bpb 4k → 8k second doubling returns 0.62× the first mbpb reduced per minute Cost versus performance 1.00 1.05 1.10 1.15 100 130 160 ctx2048 ctx4096 ctx8192 Runtime (minutes) — cost val/bpb Marginal efficiency of each doubling 0.0 1.5 3.0 2.13 2k → 4k 1.32 4k → 8k mbpb reduced per minute

The dashed line is the rate the first doubling set. The 8k run falls short of it, and the second doubling returns 0.62× the first — the shortfall is why we stopped at 4,096.

Data
RunContextRuntime (min)Validation BPB
ctx20482,048105.41.110
ctx40964,096119.51.080
ctx81928,192149.71.040
Left: validation BPB against runtime for 2k, 4k and 8k context. Right: how much loss each doubling removed per extra minute. The second doubling returns 0.62× the first.

#EXP012 · Tokenizer

Our wildly unique pretrain corpus lent us to test nanochat's built-in tokenizer variables — specifically document_cap, max_characters, and where documents start. A few findings emerged: random sampling improved compression ~1%, pushing max_characters above 80M barely improved compression, and high byte-fallback rates pointed to remaining OCR dataset junk. Despite the diminishing returns, it was cheap to train our tokenizer on our full corpus.

#EXP014 · Autoresearch

Prior to running our final model we ran 100 auto research experiments with the goal of improving our val/bpb relative to our compute budget. We adapted Karpathy's autoresearch to our model.karpathy/autoresearch — an agent edits one training file, runs a five-minute experiment, reads validation BPB, and iterates. Each experiment runs on a single H100, and the harness we ran them through is open source as bart-autoresearch. The training script runs for a fixed time budget of 5 minutes. GPT 5.6 Sol extra high was the model used for autoresearch and the loop ran for 10 hours. There were 26 improvements that reduced validation BPB from 1.244632 to 1.223533, or 1.70%. We only kept improvements to the Muon optimizer's momentum since they were the most impactful and we didn't have the time to scale up more potential improvements. Constant 0.90 was the best d12 result and was adopted for the final d32 run.

Autoresearch progress
100 experiments, 26 kept improvements, 10 hours on one H100 · baseline 1.244632 → best 1.223533 (−1.70%)
Discarded Kept Running best
1.22 1.23 1.24 1.25 1.26 1.27 0 20 40 60 80 100 baseline 1024-token training embedding LR 1.0 depth 7 RoPE base 200k SwiGLU 2.75× constant Muon 0.83 one full-context layer Experiment # Validation BPB (lower is better) 1.22 1.24 1.26 0 50 100 Experiment # Validation BPB (lower is better)

Each dot is one five-minute run; hover or tap any of them for what the agent tried. A change is kept only when it lowers validation BPB, so the running best is a staircase. The step at experiment 67 is the Muon momentum change carried into the d32.

Data
ExperimentValidation BPBChange kept
#01.244632baseline
#11.242767train with 1024-token sequences
#81.240700embedding learning rate 0.8
#91.240067embedding learning rate 1.0
#121.239271unembedding learning rate 0.003
#171.238957Muon weight decay 0.1
#281.2385957-layer transformer at width 512
#311.2385673.5x MLP expansion
#351.237959rotary base 50000
#361.237640rotary base 100000
#371.237334rotary base 200000
#431.2369322.25x parameter-matched SwiGLU MLP
#441.2366262.5x SwiGLU hidden size
#451.2350702.75x SwiGLU hidden size
#511.234631SwiGLU embedding learning rate 1.1
#531.234166SwiGLU embedding learning rate 1.05
#671.231364final Muon momentum 0.93
#681.229516final Muon momentum 0.91
#691.228225final Muon momentum 0.89
#701.227259final Muon momentum 0.87
#711.226465constant Muon momentum 0.85
#721.225522constant Muon momentum 0.83
#771.225260learning-rate warmdown over final 60 percent with low momentum
#851.223947one full-context evaluation layer
#931.223923low-momentum Muon weight decay 0.15
#941.223537low-momentum Muon weight decay 0.20
#961.223533low-momentum matrix learning rate 0.045
Every experiment the agent ran, with the running best drawn over it. The step at experiment 67 is the constant-momentum Muon change we carried into the d32.

Nanochat's original architecture is already very optimized, and a lot of the low hanging fruit have already been eaten. Many of the improvements from autoresearch were not novel. The models clearly understand the objects they manipulate at a deep level, and yet very few genuinely new ideas emerge, which makes it hard to tell if this is an artifact of the speedrun setup or a real capability of the agent models.

#Midtraining

Midtraining sits between pretraining and downstream use: instead of continuing on the same broad corpus, we shift the data mixture toward a smaller, higher-quality set of documents while decaying the learning rate. The goal is to spend the model's final optimization steps on the text we most want it to internalize, which for us meant pre-1930s math, science, technology, and medicine. This section covers both how we built the midtraining dataset, and how we scheduled the run itself.

For our midtraining dataset, we focused on finding high quality math, science, technology and medicine documents from before the 1930s. For clarity, we broke each major step into stages. We started out by extracting these documents using their subject tags from bart-dataset-v1. Then, we extracted vintage documents from the Internet Archive by filtering by topic for Math, Science, Medicine, and Technology documents using their published date. We then de-duped these documents and used an OCR filter of 0.85. To get a second opinion, we filtered out 10% of documents using additional OCR filtering with OCRoscope (0.85) and an alpha numerical ratio of 0.65. Due to math documents having many different symbols, we wanted to keep the alpha numerical ratio relatively relaxed. Then, we filtered for OCR artifacts and removed portions of the boilerplate off of documents (headers, library stamps, etc). We mimicked our filtering in our main corpus using a log prior filter (2.5, 97.5).

Finally, we made sure our midtrain dataset was as vintage as possible, running it through the footer filter and the banned words list. As we cleaned our original corpus, we learned valuable tools and methodology for dataset cleaning. We applied this knowledge to midtraining extensively. After all the cleaning in midtraining, we removed 24% of the original midtraining data leaving us with 608M tokens of high quality midtraining data — bart-midtrain.

For the midtraining itself, we mimicked HuggingFace's SmolLM with midtraining in stages of decay.SmolLM's recipe decays the learning rate in stages while progressively raising the share of high-quality data. huggingface.co/blog/smollm3 Training ran in three stages. Because midtrain documents are much shorter than bart-dataset-v3 documents, we used tokens rather than documents to measure the mixtures of midtrain to original corpus. We wanted our midtrain data to have less than 3 epochs, so we used that as a guide for the ratios at each stage.

Bart data mixture by token count
Share of tokens drawn from each corpus, across the run
bart-dataset-v3 bart-midtrain
0% 20% 40% 60% 80% 100% 0B 2B 4B 6B 8B 10B 12B 14B 16B 18B 20B 100% clean Stage 1 14B 21% 79% Stage 2 4B 46% 54% Stage 3 2B Training progress (tokens) Share of tokens 0% 25% 50% 75% 100% 0B 5B 10B 15B 20B 100% clean S1 14B 21% 79% S2 4B 46% 54% S3 2B Training progress (tokens)

20B tokens total, with stage boundaries at 70% and 90% of pretraining. Mixtures are measured in tokens rather than documents, since midtrain documents are far shorter than clean ones — the assumption that cost us 36 hours mid-run.

Data
StageTokensbart-dataset-v3bart-midtrain
Stage 114B100%0%
Stage 24B79%21%
Stage 32B54%46%
Midtrain data enters at the 70% mark and steps up again for the final 10%. Stage 2 lands at 21% midtrain, Stage 3 at 46%.

For a learning rate scheduler, we used nanochat's Warmup-Stable-Decay (WSD) starting decay at around ~35%. On the other hand, SmolLM started decaying at around ~75%.

Learning rate schedules: Bart vs. SmolLM3
Warmup-stable-decay in both cases · normalised to peak LR and to each run's own token budget
Bart — decay over final 62% SmolLM3 (11T) — decay over final 26%
0.0 0.2 0.4 0.6 0.8 1.0 0% 10% 20% 30% 40% 50% 60% 70% 80% 90% 100% decay starts 37.5% decay starts 74.5% Pretraining progress (% of total token budget) Learning rate (fraction of peak) 0.0 0.5 1.0 0% 25% 50% 75% 100% Pretraining progress

Bart spends nearly two thirds of its budget decaying; SmolLM3 spends a quarter. The shaded triangles are the decay phases: Bart's starts at 37.5%, SmolLM3's at 74.5%.

Data
RunWarmup endsDecay startsDecay over final
Bart2%37.5%62%
SmolLM3 (11T)4%74.5%26%
Decay starts at 37.5% of the token budget for Bart and 74.5% for SmolLM3 — Bart spends nearly two thirds of its run decaying.

One of the biggest pitfalls on this journey was during our final model run. We assumed that the size of midtrain documents would be roughly smaller than the original corpus. We were right, but we didn't realize how drastically smaller these midtrain documents were. Since we used document count rather than tokens to create the mixtures used for midtraining, we had to recreate our datasets with ratios with tokens rather than document counts. Since our final model was already running when we discovered this bug, we had 36 hours to create these new mixtures before our model would need them. Ultimately, we were able to get the mixtures ready for the model in time along with learning a valuable lesson — verify your intuition.

The final d32 run, either side of the pause
Training loss every 100 steps · single H100, 5 days · paused at step 5500 to swap in the rebuilt midtrain shards
Before the pause (steps 0–5500) After the fix (5500–9500) pause at step 5500
1.5 2 3 5 10 0 1k 2k 3k 4k 5k 6k 7k 8k 9k paused at step 5500 midtrain shards rebuilt Training step Training loss (log scale) 1.5 2 3 5 10 0 3k 6k 9k Training step Training loss (log scale)

The curve picks up where it left off — the seam at step 5500 is the pause spent swapping in the rebuilt shards, not a change in the run. The single low point just after it is the first logged step after the restart.

Data
SegmentStepTrain loss
Before the pause010.399
Before the pause1004.667
Before the pause2003.953
Before the pause3003.511
Before the pause4003.259
Before the pause5003.414
Before the pause6003.048
Before the pause7003.039
Before the pause8003.077
Before the pause9003.093
Before the pause1,0002.969
Before the pause1,1003.045
Before the pause1,2003.074
Before the pause1,3002.887
Before the pause1,4002.784
Before the pause1,5002.931
Before the pause1,6003.051
Before the pause1,7002.735
Before the pause1,8002.901
Before the pause1,9002.771
Before the pause2,0002.714
Before the pause2,1002.868
Before the pause2,2002.631
Before the pause2,3002.755
Before the pause2,4002.638
Before the pause2,5002.697
Before the pause2,6002.676
Before the pause2,7002.790
Before the pause2,8002.759
Before the pause2,9002.590
Before the pause3,0002.637
Before the pause3,1002.589
Before the pause3,2002.646
Before the pause3,3002.641
Before the pause3,4002.500
Before the pause3,5002.606
Before the pause3,6002.643
Before the pause3,7002.627
Before the pause3,8002.440
Before the pause3,9002.561
Before the pause4,0002.814
Before the pause4,1002.637
Before the pause4,2002.434
Before the pause4,3002.417
Before the pause4,4002.612
Before the pause4,5002.523
Before the pause4,6002.629
Before the pause4,7002.479
Before the pause4,8002.511
Before the pause4,9002.243
Before the pause5,0002.480
Before the pause5,1002.210
Before the pause5,2002.445
Before the pause5,3002.485
Before the pause5,4002.421
Before the pause5,5002.321
After the fix5,5001.578
After the fix5,6002.356
After the fix5,7002.331
After the fix5,8002.393
After the fix5,9002.451
After the fix6,0002.307
After the fix6,1002.360
After the fix6,2002.501
After the fix6,3002.277
After the fix6,4002.526
After the fix6,5002.510
After the fix6,6002.304
After the fix6,7002.295
After the fix6,8002.519
After the fix6,9002.465
After the fix7,0002.355
After the fix7,1002.358
After the fix7,2002.424
After the fix7,3002.584
After the fix7,4002.223
After the fix7,5002.408
After the fix7,6002.537
After the fix7,7002.357
After the fix7,8002.233
After the fix7,9002.304
After the fix8,0002.218
After the fix8,1002.246
After the fix8,2002.031
After the fix8,3002.219
After the fix8,4002.259
After the fix8,5002.323
After the fix8,6002.373
After the fix8,7002.108
After the fix8,8002.258
After the fix8,9002.187
After the fix9,0002.124
After the fix9,1002.021
After the fix9,2002.322
After the fix9,3002.039
After the fix9,4002.261
After the fix9,5002.309
Training loss for the final d32, either side of the pause at step 5500 where the rebuilt midtrain shards were swapped in. The run resumes on trend.

#Midtraining results

The goal of midtraining was to use the model's final optimization steps to internalize a smaller, higher-quality body of text. While we hoped this targeted phase would produce clear improvements, our evaluations do not provide conclusive evidence that midtraining was beneficial.

Per-task change in Vintage CORE from midtraining
Bart d24 · baseline 0.1369 → midtrained 0.1351
Midtrain better (Δ > 0) Midtrain worse (Δ < 0)
-0.04 -0.02 0.00 +0.02 +0.04 squad +0.031 lambada_openai +0.017 coqa +0.014 arc_challenge +0.011 hellaswag_zeroshot +0.009 hellaswag +0.006 bigbench_language_identification +0.003 jeopardy +0.003 commonsense_qa +0.002 bigbench_repeat_copy_logic +0.000 boolq -0.003 bigbench_qa_wikidata -0.004 arc_easy -0.007 winograd -0.007 bigbench_operators -0.010 piqa -0.015 agi_eval_lsat_ar -0.016 openbook_qa -0.019 copa -0.020 winogrande -0.032 -0.04 0.00 +0.04 squad +0.031 lambada +0.017 coqa +0.014 arc_challenge +0.011 hellaswag_0s +0.009 hellaswag +0.006 bb_language_identification +0.003 jeopardy +0.003 commonsense_qa +0.002 bb_repeat_copy_logic +0.000 boolq -0.003 bb_qa_wikidata -0.004 arc_easy -0.007 winograd -0.007 bb_operators -0.010 piqa -0.015 lsat_ar -0.016 openbook_qa -0.019 copa -0.020 winogrande -0.032

Nine tasks improve, ten regress, and the net is -0.002 across 20 tasks — inside the run-to-run spread measured in EXP007. Task names are abbreviated on narrow screens; hover for the full one.

Data
TaskΔ Vintage CORE
squad+0.031
lambada_openai+0.017
coqa+0.014
arc_challenge+0.011
hellaswag_zeroshot+0.009
hellaswag+0.006
bigbench_language_identification+0.003
jeopardy+0.003
commonsense_qa+0.002
bigbench_repeat_copy_logic+0.000
boolq-0.003
bigbench_qa_wikidata-0.004
arc_easy-0.007
winograd-0.007
bigbench_operators-0.010
piqa-0.015
agi_eval_lsat_ar-0.016
openbook_qa-0.019
copa-0.020
winogrande-0.032
Per-task change in Vintage CORE from midtraining on the d24. Nine tasks improve, ten regress, and the net is −0.002 — inside the run-to-run spread measured in EXP007.

On the other hand, there is no conclusive evidence that our midtraining dataset is harmful. Due to the lack of resources and time, we were not able to run ablations on different data mixtures for midtraining. Because of this, we had to rely on a meager number of experiments before scaling up our model.

#Scaling

Validation BPB against consumed tokens
The d24 pilot and the final d32 · evaluated every 250 steps · lower is better
Bart d24 — 8.8B tokens, ends 0.920 Bart d32 — 20.1B tokens, ends 0.770 midtrain stages
0.8 1 1.5 2 3 0 5B 10B 15B 20B stage 2 21% midtrain stage 3 46% midtrain Consumed training tokens Validation bits-per-byte (log scale) 0.8 1 1.5 2 3 0 10B 20B Consumed training tokens Validation bits-per-byte (log scale)

At matched input — the d24's full 8.8B-token budget — the d24 is still ahead, 0.920 against 0.960. The d32 passes it near 11.5B tokens and keeps falling to 0.770. The dashed lines are where the d32's midtraining stages begin: stage 2 at 70% of the run raises midtrain data to 21% of the mixture, stage 3 at 90% raises it to 46%.

Data
ModelStepTokensVal BPB
Bart d2400.00B3.414
Bart d242500.26B1.365
Bart d245000.52B1.195
Bart d247500.79B1.119
Bart d241,0001.05B1.108
Bart d241,2501.31B1.097
Bart d241,5001.57B1.084
Bart d241,7501.84B1.066
Bart d242,0002.10B1.078
Bart d242,2502.36B1.070
Bart d242,5002.62B1.053
Bart d242,7502.88B1.041
Bart d243,0003.15B1.043
Bart d243,2503.41B1.032
Bart d243,5003.67B1.020
Bart d243,7503.93B1.012
Bart d244,0004.19B1.005
Bart d244,2504.46B0.999
Bart d244,5004.72B0.973
Bart d244,7504.98B0.973
Bart d245,0005.24B0.968
Bart d245,2505.51B0.969
Bart d245,5005.77B0.962
Bart d245,7506.03B0.960
Bart d246,0006.29B0.955
Bart d246,2506.55B0.950
Bart d246,5006.82B0.946
Bart d246,7507.08B0.941
Bart d247,0007.34B0.937
Bart d247,2507.60B0.934
Bart d247,5007.86B0.932
Bart d247,7508.13B0.927
Bart d248,0008.39B0.924
Bart d248,2508.65B0.922
Bart d248,3528.76B0.920
Bart d3200.00B3.414
Bart d322500.52B1.249
Bart d325001.05B1.148
Bart d327501.57B1.083
Bart d321,0002.10B1.057
Bart d321,2502.62B1.025
Bart d321,5003.15B1.015
Bart d321,7503.67B1.004
Bart d322,0004.19B0.993
Bart d322,2504.72B0.963
Bart d322,5005.24B0.974
Bart d322,7505.77B0.973
Bart d323,0006.29B0.974
Bart d323,2506.82B0.975
Bart d323,5007.34B0.968
Bart d323,7507.86B0.969
Bart d324,0008.39B0.959
Bart d324,2508.91B0.960
Bart d324,5009.44B0.957
Bart d324,7509.96B0.950
Bart d325,00010.49B0.936
Bart d325,25011.01B0.935
Bart d325,50011.53B0.900
Bart d325,75012.06B0.903
Bart d326,00012.58B0.888
Bart d326,25013.11B0.895
Bart d326,50013.63B0.888
Bart d326,75014.16B0.882
Bart d327,00014.68B0.878
Bart d327,25015.20B0.868
Bart d327,50015.73B0.871
Bart d327,75016.25B0.866
Bart d328,00016.78B0.843
Bart d328,25017.30B0.834
Bart d328,50017.83B0.832
Bart d328,75018.35B0.807
Bart d329,00018.87B0.804
Bart d329,25019.40B0.772
Bart d329,50019.92B0.770
Bart d329,60020.13B0.771
Validation bits-per-byte against consumed tokens for the d24 and the final d32. At the d24's full budget the smaller model is still ahead; the d32 passes it near 11.5B tokens and ends at 0.770.

Scaling up to the d32, we knew we needed a hardware upgrade. Previously, we were using A100s with Flash Attention 2 for all experiments. We set our eyes on H100s for their 2–3 times speed up with Flash Attention 3. After researching multiple different options for leasing, vast.ai was by far the most price efficient.vast.ai — a marketplace for spot GPU rental, priced per host rather than per region. We considered running a 4x or 8x cluster to speed up the runs. Preparing for the 4x and 8x clusters, we created a docker image that has all of our dependencies and saves us 30 minutes on setup. We ran our first d24 on a 4x cluster to test the new docker image. One difficulty in using a multi-GPU cluster is that resuming requires similar hardware unless you are willing to drop the momentum of the optimizer, losing around 20 steps of learning. While there are possible solutions to circumvent this problem, we decided to resume our runs on the same hardware to spend our time in more valuable areas.

Compared to other options for leasing, vast.ai's GPUs shined brightest when using single devices. Since there are far more options to select from for single devices, it was far easier to find a bargain — especially past 11pm EST. While typical H100s cost ~$2–3 an hour, we were able to snag one for 99 cents an hour. Starting the d32 run Friday morning and with our budget approaching its maximum, we decided to opt for the single H100 as speed was not a huge priority. The run took 5 days on a single H100. Throughout our run we held a 60 MFU on our H100, our loss had minimal spikes and held a consistent 50k tokens per second. We had to pause the run at step 5500 (57%) to fix the midtraining shards, this resulted in 1 hour and 30 min of down time. The run remained healthy throughout, and the full d32 run is public on Weights & Biases, as are the two d24 runs that preceded it.The d24 ran in two parts on the 4x cluster: e69c1e59 · 2fa1b744

#Post-training

Making our model "feel" and act like an individual from the 1930's falls under the category of supervised fine tuning. SFT datasets are abundant for normal models, but for models trained exclusively on pre-1930's data the vast range of options vanish.A few do exist — vintage-ft-v2 and Violet-sft among them — but none at the scale or specificity we needed. Because authenticity was a priority for us it seemed smart to start by sourcing authentic question and answer (q/a) content that's already available online from the time period. This led to catechisms — texts specifically written in a question and answer format. About thirty catechisms total were sourced with topics such as engineering, logic, music, world history, and philosophy. A thorough extraction process included filtering, cleaning, enriching, and joining the q/a pairs. About 12k rows (with ~25% being multiturn conversations) were collected (authentic-pre1930-sft). For comparison, Hla's GPT-1900 model was trained on roughly 50k rows.Vintage-exam-qa. See michaelhla/gpt1900.

Research suggests that more important than size is quality and diversity.Zhou et al. LIMA: Less Is More for Alignment, 2023. arXiv:2305.11206 Since authentic sources would never be able to cover the many types of behaviors we wanted our model to exhibit, we pivoted to generating q/a pairs instead. Inspired by TypeWriterLM's History-SelfInstruct pipeline, we opted for grounded-answer based generation. In essence, pull excerpts from our corpus, trim and shape them to make good answers, and then generate a feasible question based on the given answer. This means generation and anachronism is halved, only the questions can carry an anachronistic risk. Further, because models typically mask the questions when fine tuning, in theory, the model never updates its weights directly on anachronistic text (though context still picks it up). To encourage diversity, about ten categories of q/a pairs were drafted consisting of narrative generation, knowledge recall, opinions, STEM reasoning, refusal calibration, etc.

Total counts per question category

synthetic-pre1930-sft · 416k rows, ~70M tokens

multiturn_qa131.8Kknowledge_qa65.9Kcomposition_qa40.7Kopinion_qa38.6Knarrative_grounded38Knarrative_fiction31.7Kreasoning_qa24.1Kverse_qa17.4Khow_to_qa17Kstem_reasoning8.3Kcalibration_qa2.7K
Data
GroupRows
multiturn_qa131.8K
knowledge_qa65.9K
composition_qa40.7K
opinion_qa38.6K
narrative_grounded38K
narrative_fiction31.7K
reasoning_qa24.1K
verse_qa17.4K
how_to_qa17K
stem_reasoning8.3K
calibration_qa2.7K
Roughly half the dataset is knowledge and multiturn recall. The long tail — STEM reasoning at 2.0% and refusal calibration at 0.7% — is small by design but carries the behaviours we most wanted.

Excerpts were algorithmically cherry-picked from the corpus, and then prompts were refined for each category to avoid context-bare, irrelevant questions, and phrase-repetition.

Examples of bare questions:

  • "What error does the writer discover in the note from the Insurance Society concerning the premium?" — references the passage's author and an unnamed note; unanswerable to anyone who didn't read it.
  • "What was the nationality and age of the captain of the corsair?" — which corsair?

Questions after tuning the prompt:

  • "What was the method of cure, as given in 1857, for the disease in cattle evidenced by a frequent slight kicking of the hind foot?"
  • "…the steamer built by Workman & Clark, of Belfast, which possessed a greater carrying capacity than…"

The result is synthetic-pre1930-sft, a monstrous 416k row (~70M tokens) dataset consisting of about 33% multiturn conversations, and roughly half of the whole dataset purely knowledge-based q/a pairs. The remaining half consists of the less common question types meant to encourage diverse behavior and reduce hallucinations. Two highlights: STEM reasoning, a modestly sized intellectual reasoning set, and refusal calibration, a vital non-answer set fostering honesty. Pairs are categorized, dated, and graded making this highly adaptable to different cutoff dates, SFT curriculums, and model personalities.

Grade distribution within each category
synthetic-pre1930-sft · share of pairs in each grade band
95–100 90–94 80–89 65–79 40–64 0–39
0% 20% 40% 60% 80% 100% knowledge_qa calibration_qa composition_qa narrative_grounded verse_qa reasoning_qa opinion_qa multiturn_qa narrative_fiction how_to_qa stem_reasoning 0% 50% 100% knowledge_qa calibration_qa composition_qa narrative_grounded verse_qa reasoning_qa opinion_qa multiturn_qa narrative_fiction how_to_qa stem_reasoning

Shares are read from the grading run's own figure and are accurate to about a percentage point. Grading is subtractive: a pair starts at 100 and loses points for each defect the grader finds.

Data
Category95–10090–9480–8965–7940–640–39
knowledge_qa80.5%13.7%2.6%0.7%2.5%0.0%
calibration_qa77.3%10.9%3.3%1.1%7.4%0.0%
composition_qa58.0%20.8%9.0%3.5%8.7%0.0%
narrative_grounded47.5%25.5%11.5%3.0%12.5%0.0%
verse_qa43.8%23.2%12.2%3.6%17.2%0.0%
reasoning_qa42.0%27.5%16.5%6.0%8.0%0.0%
opinion_qa40.2%33.3%18.8%4.5%3.2%0.0%
multiturn_qa40.0%28.3%16.5%5.2%10.0%0.0%
narrative_fiction24.6%15.4%9.4%3.4%47.2%0.0%
how_to_qa23.5%28.7%26.3%14.0%6.8%0.7%
stem_reasoning23.2%26.3%26.3%13.4%10.2%0.6%
Grade distribution within each category. Knowledge and calibration questions grade highest; fiction and how-to prompts are where the generator struggles most.

Grading is designed to answer the question, "How well does the q/a pair teach the model the intended goal of its category?" As shown in Figure 21, grading dramatically changes the landscape of the dataset. Around 196k q/a pairs are graded ≥ 95, and this increases to roughly 297k (+100.6k) for pairs ≥ 90. Grading is subtractive: if the model finds defects in the q/a pair then it loses points. Differences in scores above 95 can, therefore, reasonably be reduced to noise. As such, we found that questions graded below 80 were not worth including in our final fine tuning spread.

Examples of graded pairs
synthetic-pre1930-sft · hover a pair for its score and the deductionone pair from each band the grader used
95–100 90–94 80–89 65–79 40–64
  1. QWhat was the cause of Décazes's anxiety?

    AI do not know and cannot conjecture the cause of Décazes's anxiety.

    97 · graded 95–100 · calibration_qa — asked for a cause the passage never gives, and it says so outright.

  2. QWhat is the law of flotation for a body floating in a liquid?

    AA floating body displaces its own weight of the liquid in which it floats.

    91 · graded 90–94 · knowledge_qa — exact and self-contained, but it answers in the question's own words rather than adding to them.

  3. QOught a man to trace his ancestry to apes or to Adam?

    AThey trace their ancestry to apes- To Adam I trace mine.

    85 · graded 80–89 · opinion_qa — a couplet that takes a side, but the run-in dash is OCR damage, not punctuation.

  4. QHow is transparent varnish made for wood?

    AIn 1 quart of alcohol mix 2 oz. gum mastick, 4 oz. of Venice turpentine, and 8 oz. of gum juniper.

    75 · graded 65–79 · how_to_qa — quantities but no method: a how-to that never says how.

  5. QIs it just that the people are cheated by political humbugs?

    AThus poor mankind are cheated, By humbugs of the day;

    42 · graded 40–64 · opinion_qa — flagged “answer appears to end mid-sentence or mid-structure”; the verse stops on a semicolon.

The same rubric across its range, one pair to a band. Each mark carries its score, its category, and what the deduction was for.

Speaking of fine tuning spreads: Chen et al. and Kaur et al. demonstrate that a small, rich dataset consistently outperforms a much larger, noisy one.Chen et al. AlpaGasus: Training a Better Alpaca with Fewer Data, 2023. arXiv:2307.08701 · Kaur et al. Instruct-SkillMix, 2024. arXiv:2408.14774 Applying these insights to our dataset means that we shouldn't naively dump all of the q/a pairs into our model. A fine tuning "curriculum" — as in a specific spread of learning material — needed to be concocted that would suit our model best. Six curricula were created to explore different ways of distributing the data. We wanted to find what drove the best model performance: data quantity, data quality, token distribution, domain scheduling, or domain difficulty. Surprisingly, we found that perhaps less isn't more in our case: our model chatted significantly better on the largest quantity of data. A cumulative 362k rows (~75M tokens) across 3 stages, progressing from easier to harder questions, thresholded at grades above 80, led to the best chatting experience by far. The 362k training rows were even later pushed to ~650k in a further iteration.

#Dataset limitations

Right at the end of the project, significant limitations surfaced regarding our fine tuning coverage. Experimentation on the smaller d24 model displayed weird quirks. For example, the model could handle questions just fine, but ask the very same question without a question mark and the model would collapse down to its base model behavior, spewing hundreds of incoherent tokens and often degrading into repetition. Similarly, the model couldn't handle unparseable text, typos, or simple conversational openers such as "Hello", or "How are you?". These types of inputs weren't included in the fine tuning dataset, so the model had no register for them.

The same question, one character apart

What is lunch?

A lunch is a large dinner consisting of several courses served together. One lunch consists of plums and cherries served in paper bags, another of oranges and nuts in paper cases. One lunch consisted of nuts, oranges and mushrooms in paper cases. One lunch consisted of mushrooms and nuts.

289 characters · 49 words · ends

What is lunch

A dinner in which food is served and served to guests while they eat. It is the practice in the United States. In this country, the guests sit down and eat as they choose. At the end of the meal, guests are served, or served, in the same way as at a dinner. A long time ago, a dinner, or lunch, was a meal, while now it is a term employed to describe a meal served to guests. In the first, there is confusion; in the second, order. In this country, the waiters wear a distinctive uniform. They are uniformed in England or Germany, or uniformed according to English regulations. They are, however, not uniformed in this country. They wear the brass button of their coats, a button they claim to be able to make; and they are not uniformed. The American waiter is not uniformed. He may be a good manager, and he may be a bad one. He is apt to be a big talker, and he is apt to be a blusterer. The American waiter, to tell the truth, is no more apt to be either than is the American barkeeper. It is to be hoped that as barkeeper, or what amounts to the same thing, as waiter, he will get discharged. In England and Germany, the waiters, generally, are uniformed; and they wear a distinctive uniform, the Prussian. In England and Germany, waiters have their distinctive uniforms, and the English and German waiters wear theirs, as well as that of the waiter. In the first, they have the English dress of the waitress; in the second, they wear a dress distinctive of the barkeeper. The American waiter is a uniformed man, and he is not a waiter, but he is part of the regular service, and the regular service does not take the place of the waiter, but it helps to create the waiter. It is a mistake to call the waiter by the name of "soldier." The waiter is a man, part of the camp, and part of the town. He has a name of his own, and is not a man, but a man belonging to the army, the navy, the police, the post-office, the railroad, the street-car, the store, the bank. He is a boy belonging, not to this place, but to that; and it is a mistake to call him "soldier." The American waiter is not a soldier, but a regular.

2,118 characters · 409 words · still going

The same question asked twice, the second time without its question mark. The reply runs seven times longer and never finds its way back to lunch.

The final d32 (2.8B) model appeared to generalize much better and was able to handle out-of-distribution inputs well. Still, the issues persisted and it led to a few last minute techniques being applied. One solution was to introduce artificial noise into the existing SFT corpus. Periodically and deterministically removing words, transposing letters, and adding or removing capital letters, punctuation, and spacing from questions made the model more sturdy at inference.

Larger-scale input corruption and diversity still needed to be addressed, though. So, a second solution was to create a complementary dataset aimed at teaching the model to respond as intended to outlying data. We made Robust-SFT for this purpose. User input that was uninterpretable, such as heavy typos, gibberish, or completely random character strings were paired with corpus-verbatim non-answers, such as "I beg your pardon" or "That is beyond me". 2,000 rows were synthesized across 11 classes of input corruption. Around 200 period and era grounding q/a pairs were synthesized to give the model some stiffness in this area. This comes with the highest anachronism risk, however, as they aren't corpus-grounded. Additionally, two types of conversational data were added: one which handled confusing but grammatically correct statements, and another prioritizing multiturn conversations and exchanges, which was very rare in the original dataset. 7,338 rows total, distributed across each stage of the original SFT dataset.

Despite our efforts, the issues weren’t completely fixed. Normal, structurally correct input without punctuation was still very volatile (pointing to an undertrained and underrepresentative fine tune). With time running low, we couldn’t spend more time curating data for the model to train on. We had to include a hotfix. The simplest change was to artificially append punctuation to the end of the user’s input. When experimenting, this cut the deterioration rates in half. Even better, though, was to prime the chat by prepending one full conversation turn. It’s like system prompting but simpler. The model essentially has the context of an earlier pre-specified introduction that the user doesn’t see. Our model is sufficiently trained on multiturn conversations, so it “recognizes” every interaction and responds properly even to those without proper grammar. The model that came out of all this — the C3 curriculum plus the robustness rows, branched from the base d32 at step 9,600 — is bart-sft, and it is what the demo serves.

#Evaluations

ChatCOREnanochat's chat evaluation: scripts/chat_eval.py · tasks/ is another benchmark suite reserved by Karpathy for models that come through the nanochat pipeline. The synthetic-pre1930-sft dataset was created with ChatCORE in mind. Unfortunately, evaluations on this benchmark weren’t run during development. It turns out that ChatCore is quite a biased benchmark and is mostly in place just to see how well a model learns Karpathy’s built in fine tuning tasks. In fact, ChatCORE is the reserved, set-aside validation set of nanochat’s fine tuning tasks. So when we don’t fine tune our model on his examples, the model does… poorly:

ChatCORE, the same base fine-tuned two ways
Bart d32 · nanochat's own eval suite · each ring is one task, filled to its score. ChatCORE is centred so that 0 is chance and 1 is a perfect score.
ARC-Easy ARC-Challenge MMLU GSM8K HumanEval chance
0.135 ChatCORE nanochat default its own SFT data, size-matched −0.011 ChatCORE pre-1930 curriculum C3 + robustness, our SFT data
Data
Task nanochat default pre-1930 curriculum Chance
ARC-Easy45.5%24.6%25%
ARC-Challenge37.9%23.0%25%
MMLU34.7%23.1%25%
GSM8K6.4%0.0%0%
HumanEval3.7%0.0%0%
ChatCORE0.135−0.0110.000
One d32 base checkpoint, fine-tuned twice: on nanochat's own SFT data, and on our pre-1930 curriculum. Measured on nanochat's suite, the vintage model sits at chance.

Because of the ChatCore oversight, Bart's post training remains largely unevaluated.

#Hosting

With our scaling strategy in place, the next challenge was deciding where and how to actually run our models. Hosting GPU-backed inference presented its own set of constraints, chief among them cost. From the outset, we sought a serverless GPU solution, evaluating Modal, RunPod, and Beam among several other providers.modal.com · runpod.io · beam.cloud Since our budget could not accommodate keeping a GPU running around the clock, we prioritized fast cold starts and automatic scaling. After weighing our options, we settled on Beam for a number of reasons. Foremost among them was the generosity of Beam's free tier, which offers up to five concurrent GPUs, free credits, automatic scaling, and access to affordable consumer hardware such as the RTX 4090 at just $0.69 per hour. The performance exceeded our expectations as inference was effectively instantaneous and cold starts were averaging only ten seconds.

#Finances

Training a language model is constrained as much by compute costs as by technical considerations. For this project, we worked with the limitations of being students without external funding or institutional support. Despite these constraints, we were able to train and evaluate a model that we are proud of, while keeping our overall expenses relatively modest.

For our ablation experiments, our primary expense was Google Colab Pro, followed by compute rented through Vast.ai. We used Hugging Face Pro for hosting our datasets and model checkpoints. For model development and experimentation, we used Claude Code alongside Codex, while DeepSeek and OpenCode were used primarily for SFT development and generation. Finally, for our larger d24 and d32 runs, we relied on Vast.ai due to the availability of comparatively inexpensive GPUs.

In total, our recorded expenses were approximately $800. This was not a fixed budget established at the beginning of the project; rather, it represents what we ultimately spent across development, training, and experimentation. The breakdown of these costs is shown below:

ServiceCost
Vast.ai — GPU leasing$227.18
DeepSeek API$174.93
Google Colab$70.00
Claude Code$140.00
Codex$30.00
Hugging Face Pro$30.00
OpenCode$10.00
Beamfree tier
Total$682.11

#Conclusions

Our goal throughout this project was to determine how far careful data selection, training decisions, and limited compute could take a vintage language model. Exploring this niche field gave us insight into countless problems that the frontier labs do not experience, such as the lack of large amounts of high quality SFT data. The results are not uniformly conclusive, but they give us a useful account of what worked, what did not, and — perhaps most importantly — what we would change in a future training run.

Looking back, there are several aspects of the training process that we would approach differently. During pretraining, we would have spent more time refining and validating the dataset before moving on to ablations. We also would have conducted a more thorough set of learning-rate experiments rather than relying primarily on established recommendations. For midtraining, we would have applied more aggressive filtering to the corpus, potentially using an LLM-based quality filter similar to the approach used by FineWeb-Edu. Additionally, budget and time constraints prevented us from systematically testing different mixtures of midtraining data, leaving an important part of the training recipe unexplored. Finally, delays in producing our final model compressed the timeline available for post-training, limiting the amount of time we could devote to refining the SFT dataset.

Despite these shortcomings, the process was an invaluable learning experience. Each limitation exposed a different challenge in training language models, from dataset construction and hyperparameter selection to experiment design and project management. While there are undoubtedly many things we would change in hindsight, these lessons provide a much stronger foundation for approaching our next training run.

Despite these challenges, we are proud of how much we accomplished over the summer while balancing the project with other commitments, including summer coursework. As researchers, we learned an enormous amount about training and evaluating language models, and we are especially proud of the resources we produced. bart-dataset-v3 is one of the first open-source filtering efforts focused on Institutional Books textbooks, while Vintage CORE provides a valuable evaluation benchmark for future vintage models. Although our SFT dataset is not perfect, the work required to create it demonstrated just how challenging building specialized SFT datasets can be. Most importantly, we are proud of how we worked together through deadlines, problems, setbacks, and successes. We are proud of the model we produced and what we learned along the way. As the first project for our lab, we consider Bart an enormous success and a valuable foundation for our future work.

#Acknowledgements

This project is built on work others did in the open.

#Citation

Please cite this work as:

@article{unboundedlabs2026bart,
title   = {Bart: A Vintage LLM},
author  = {Unbounded Labs},
year    = {2026},
journal = {Unbounded Labs Research},
note    = {https://unboundedlabs.ai/blog/bart}
}