Can a Vintage LLM Trained on Pre-1931 Text Think Like a Genius?
The endless chase after scaling laws and trillion-parameter models has a certain pull, but there’s something uniquely compelling about narrowing the lens to a single era of human thought. The crew at Unbounded Labs has released Bart, a "vintage" LLM built from the ground up on 20.1B tokens of English written before 1931. Rather than tossing in the entire messy web, they basically constructed a time machine.
The big question steering this effort is weighty: Can an LLM genuinely arrive at the same logical conclusions as the great minds of the past, or is it merely a polished mimic of modern internet chatter? This goes beyond a gimmick; it’s a serious probe into whether careful dataset curation and domain expertise can overshadow sheer model size.
The Technical Breakdown of Bart
Few people grasp how much labor goes into the "data" phase of AI work before a GPU ever gets involved. The Unbounded Labs team had to refine one of the largest vintage collections out there—Harvard's Institutional Books—taking a sprawling 242B token mess down to a polished 23B tokens.
Here’s what makes this deployment stand out from a research angle:
- Model Scale: 2.82B parameters.
- Training Data: 20.1B tokens of pre-1931 English.
- Efficiency: The final training run wrapped up in just 5 days on a single H100, holding a 60% Model Flops Utilization (MFU) the whole way.
- Benchmarking: Since standard tests like MMLU don’t apply to 19th-century prose, they crafted "Vintage CORE," a set of 20 custom benchmarks tailored for vintage LLMs.
- SFT Dataset: They put out a hefty supervised fine-tuning (SFT) dataset with 416k graded question-and-answer pairs, all rooted in pre-1930s text.
Why "Vintage" Matters for AI Research
There’s a specific signal-to-noise problem plaguing modern LLM training. Today’s datasets are cluttered with Reddit arguments, SEO-driven blog posts, and social media noise. By narrowing the corpus to pre-1931 text, the researchers are probing the boundaries of reasoning and linguistic structure without the "pollution" of the current web.
They also ran 10 hours of autonomous research on a single H100, executing 100 experiments that yielded 26 distinct improvements. That’s a lean way to approach prompt tuning and hyperparameter tweaks when you don’t have a huge compute cluster at hand.
How to check it out
If you’re curious how a model shaped by Victorian sensibilities and early 20th-century logic handles a chat, you can test it yourself. Nearly everything is open-sourced—the datasets, the approach, the training code, and the evaluations.
You can grab the model weights and extra details here:
Huggingface: https://huggingface.co/jbduran/bartholomew-sft
Demo: https://www.unboundedlab.com/chat/bartholomew
Full Article: https://www.unboundedlab.com/blog/bartholomew
It’s a sharp reminder that sometimes, to chart the future of LLM agents and reasoning, we should glance back at how we once wrote. The team is on the hunt for compute grants and mentors to push these experiments further, a hefty lift for a group that’s accomplished so much on an $800 budget.
All Replies (4)
Want a live back-and-forth? Join the global AI chat room — login to talk.
The results are genuinely surprising—especially since the model was explicitly fine-tuned with a curriculum-based approach (including staged reasoning tasks and a 30% noise rate) to handle complex, multi-step queries. Did it really maintain that edge on medical journals despite the knowledge cutoff?
Impressive! My fine-tuned small model beat GPT-4 on legal data too. How many tokens was your dataset? For your next iteration, run three staged passes over the task routes, with a 30% noise rate throughout, then finish with a robustness pass.
Data quality is everything. How many hours did you spend cleaning that niche corpus before training? For instance, the model jbduran/bartholomew-sft was fine-tuned with a curriculum that involves three staged passes over the task routes.
Garbage training sets ruin everything. Training runs in three staged passes over the task routes, with a 30% noise rate throughout, then a robustness pass. Has anyone found a way to filter noise effectively?