Talkie-1930: Pre-1931 Giant AI Model
Tech Talkie-1930: Pre-1931 Giant AI Model An artificial intelligence born from the dusty pages of history is attracting attention by wiping away all the dirt from modern benchmarks. The 13 billion parameter open-weight language model named Talkie-1930 was trained on 260 billion tokens of text published before January 1, 1931. Public domain sources such as books, newspapers, scientific journals, patents, and court records were used. This strict cutoff date prevents test data from leaking into the training set from the outset and makes AI generalization studies flawless. The model, powered by Claude Sonnet 4.6, is publicly accessible at talkie-lm.com/chat. Talkie-1930s Unique Training Data The non-profit team led by Nick Levine, David Duvenaud, and Alec Radford developed the model using Anthropics computing power. The dataset is entirely web-free and public domain: 260 billion tokens primarily from books and newspapers before 1931. This approach eliminates the pollution created by internet data. The cutoff date prevents train-test contamination in modern benchmarks, providing a pure generalization test. Models Technical Specifications and PerformanceFeatureValueNumber of Parameters13 BillionTraining Tokens260 BillionCutoff DateJanuary 1, 1931LicenseApache 2.0CheckpointsBase (auto-completion) + Chat (instruction-tuned) Two checkpoints are available on Hugging Face. In technical analysis, the model shows peak generalization in 1950s-60s events; abstraction power is measured without data freshness. Responses