GPAI Ledger › Inkling Small (Thinking Machines) › Capture 11 Aug 2026
Inkling Small — capture 20260811T103840Z
| Provider | Thinking Machines |
|---|---|
| Target | provider site — https://thinkingmachines.ai/training-data-documentation/ |
| Fetched (UTC) | 2026-08-11T10:38:40Z |
| Stored file | a2bfa962f8ed746dce7080d9fca841dcb24ac7ad2b10cb6a8c5ac30459355770.html.txt (19,868 bytes) (served with a .txt suffix so the captured page cannot run scripts on this site; bytes are identical — the SHA-256 verifies against this file) |
| SHA-256 | a2bfa962f8ed746dce7080d9fca841dcb24ac7ad2b10cb6a8c5ac30459355770 |
| OpenTimestamps proof | a2bfa962f8ed746dce7080d9fca841dcb24ac7ad2b10cb6a8c5ac30459355770.html.ots (calendar-attested; anchored in bitcoin over time) |
| Wayback | Wayback snapshot, 2026-07-26 01:47 UTC (pre-existing snapshot returned by the Wayback Machine — witnesses the page before this capture) |
| Prior capture of this target | — first capture of this target |
| Notes | text_sha256 recorded 20 Aug 2026 from the extracted.txt stored at capture time (bootstrap captures predate this field); raw bytes unchanged |
Verify: sha256sum a2bfa962f8ed746dce7080d9fca841dcb24ac7ad2b10cb6a8c5ac30459355770.html.txt must equal the hash above (the filename IS the expected hash); ots verify a2bfa962f8ed746dce7080d9fca841dcb24ac7ad2b10cb6a8c5ac30459355770.html.ots -f a2bfa962f8ed746dce7080d9fca841dcb24ac7ad2b10cb6a8c5ac30459355770.html.txt (opentimestamps.org) proves the capture time (fresh proofs report 'pending' until bitcoin-anchored, typically within a day).
Extracted text
Machine-extracted text (layout may be lost; the authoritative content is the stored file above).
Training Data Documentation - Thinking Machines Lab THINKING MACHINES Inkling Tinker Connectionism News Join us Home Inkling Tinker Connectionism News Join us Training Data Documentation This document provides information regarding the datasets used by Thinking Machines Lab to develop its generative artificial intelligence systems and services, including its models (collectively “AI services”). This information reflects general practices, not specific to any single model. 1. Dataset Sources Thinking Machines Lab’s AI services are developed using a variety of data sources, including publicly available data, data provided by partners via access agreements, and internally generated data, including synthetic data. 2. Intended Purpose Thinking Machines Lab uses the collected data to develop its AI services, which are designed to understand and generate content, which may include text, images, audio, and video. The datasets are selected and curated to enable the AI services to develop broad capabilities, which may include language understanding, reasoning, visual comprehension, and audio and video processing. 3. Data Types and Amount Thinking Machines Lab’s AI services are developed using datasets containing trillions of tokens. The training data consists of a broad variety of content types, including text, images, audio, and video, drawn from publicly available sources or acquired from third-parties. Thinking Machines Lab also uses synthetically generated data to supplement its training datasets. 4. Inclusion of Public Domain or IP-protected Data The datasets Thinking Machines Lab uses to develop its AI services includes content that is in the public domain as well as content that may be subject to intellectual property protection. 5. Data Acquisition Thinking Machines Lab’s services were developed using publicly available content obtained from the open internet and publicly accessible data repositories. Certain datasets were also obtained from third parties. 6. Inclusion of Personal Information or Aggregate Consumer Information Thinking Machines Lab’s training datasets may contain personal information or aggregate personal information, as defined under applicable laws, such as names or other details that individuals share publicly on the internet or with Thinking Machines Labs. 7. Cleaning, Processing, and Other Modification to Datasets Thinking Machines Lab takes steps to clean, process, and modify datasets used for its AI services. These processing steps, which vary by data type, may include deduplication and filtering to remove junk or other low-quality data, and are intended to improve the data’s usefulness for model training and support the responsible development of the AI services. 8. Data Collection and Use Timeline Thinking Machines Lab collects and uses datasets to train and improve its AI services on an ongoing basis. Thinking Machines Lab began collecting and using datasets for model development in 2025. 9. Use of Synthetic Data Thinking Machines Lab generates and uses synthetic data in the development of its AI services. Model Summary: Inkling Inkling-Small Models Inkling Inkling-Small Interaction Tinker Overview Docs Company About Connectionism News Careers Legal Terms of service Privacy notice Security Disclosure policy Connect X LinkedIn YouTube Thinking Machines Lab © 2026