GPAI Ledger › Anthropic trust-center bundle (Anthropic) › Capture 2 Sep 2026
Trust-center document bundle — capture 20260902T061659Z
Filed under Anthropic — Anthropic trust-center bundle, the source this project captured it from; the document itself is the filing of the model named above.
| Provider | Anthropic |
|---|---|
| Target | provider site — https://trust.anthropic.com/doc/trust-zip?r=… (signed URL; token masked, not linked) |
| Fetched (UTC) | 2026-09-02T06:16:56Z |
| Stored file | 16a06978e951ed6ed4b95d19b1e5cba180505a23efef21c63713b2c7e6017925.zip (1,552,029 bytes) |
| SHA-256 | 16a06978e951ed6ed4b95d19b1e5cba180505a23efef21c63713b2c7e6017925 |
| OpenTimestamps proof | 16a06978e951ed6ed4b95d19b1e5cba180505a23efef21c63713b2c7e6017925.zip.20260902T061659Z.ots (calendar-attested; anchored in bitcoin over time) |
| Wayback | not saved |
| Prior capture of this target | 4d2f4ba23f6af556bd0fad6cbc2fa40c1fad17fc98d40f4834d06ec32bd3aa14 (captured 2026-08-19T20:25:31Z) |
| Notes | downloaded via the bulk-download link publicly offered on the provider's trust-center page (the link embeds a rotating token) |
Verify: sha256sum 16a06978e951ed6ed4b95d19b1e5cba180505a23efef21c63713b2c7e6017925.zip must equal the hash above (the filename IS the expected hash); ots verify 16a06978e951ed6ed4b95d19b1e5cba180505a23efef21c63713b2c7e6017925.zip.20260902T061659Z.ots -f 16a06978e951ed6ed4b95d19b1e5cba180505a23efef21c63713b2c7e6017925.zip (opentimestamps.org) proves the bytes existed no later than the attestation time — an upper bound on the capture time; the fetch time above is the archive's own record (a freshly captured proof reports 'pending' here: the calendars anchor within hours, but this archive only upgrades the stored proof to its anchor on a later run, so expect a day or two). ots verify needs a local Bitcoin Core node (a pruned one is fine); without one, ots info on the proof prints the attesting block height and merkle path to check on any block explorer.
Files in this bundle
| File | SHA-256 | |
|---|---|---|
| Claude Mythos 5 and Claude Fable 5 Training Data Summary .pdf | 765afb9d4e62fa40… | Art. 53 summary |
| Claude Opus 4.7 Training Data Summary .pdf | 184c573a50ec4e70… | Art. 53 summary |
| Claude Opus 4.8 Training Data Summary .pdf | ccac79227fc4e9fd… | Art. 53 summary |
| Claude Opus 5 Training Data Summary .pdf | 06dfc855a9905ffd… | Art. 53 summary |
| Claude Sonnet 5 Training Data Summary .pdf | c84c83a1215c56cf… | Art. 53 summary |
| Claude Training Data Summary (Fable 5.1 + Mythos 5.1) PDF.pdf | ff3ed62ab5d93f58… | Art. 53 summary |
| _Claude Mythos Preview Training Data Summary .pdf | 038c305c4760c4d2… | Art. 53 summary |
Extracted text (Art. 53 summaries)
Machine-extracted text (layout may be lost; the authoritative content is the stored file above).
Files recorded but not stored
These bundle members are outside Art. 53 scope (some carry confidentiality markings); they are recorded by name and SHA-256 so their identity stays provable, but their bytes are not archived or served.
| File | SHA-256 |
|---|---|
| AB 2013 Training Data Documentation [June 2026].pdf | df55f60c9bc1cd11… |
| ACR for Claude Android Enterprise - May 2026 - Anthropic PBC.pdf | 541652aea109b82d… |
| ACR for Claude Web Enterprise - April 2026 - Anthropic PBC.pdf | b428badebb1568f7… |
| ACR for Claude iOS Enterprise - May 2026 - Anthropic PBC.pdf | 35f9ea604acc09f8… |
| Anthropic Statement on Modern Slavery Act 2015.pdf | d1d71f9119098a4e… |
| CVE-2026-22561 - DLL Search Order Hijacking in Claude for Windows installer.pdf | 50b69a0273d34256… |
| Claude Desktop 3P Security Overview.pdf | a3fcdcc61d5802de… |
| Claude Mythos 5 + Fable 5 Model Documentation Form v2.pdf | d4c04c6e2beb0f58… |
| Claude Mythos Preview Model Documentation Form v2 [PDF].pdf | 9fb3d243e0955890… |
| Claude Opus 4.7 Model Documentation Form for Downstream Providers_v2.pdf | 2bebf9186e1b22fd… |
| Claude Opus 4.8 Model Documentation Form for downstream providers_v2 (1).pdf | 6645a5afcaad0eb0… |
| Claude Opus 5_ Model Documentation Form for downstream providers [PDF].pdf | 5a740095beed9ef0… |
| Claude Sonnet 5 Model Documentation Form v2.pdf | df248a415f6f2b03… |
| Claude in Excel & PowerPoint Architecture Overview.pdf | 8dc1b3025d9368e5… |
| Frontier Compliance Framework_July 2026.pdf | 8e4d91e12861218e… |
| Model Documentation Form [Fable 5.1 & Mythos 5.1] for Downstream Providers PDF.pdf | 79fa36acac8578b8… |
| Office Agents 3P Architecture.pdf | 21214bde3519c4d3… |
| [Anthropic Ireland Limited] Cyber Essentials Certificate (2025).pdf | 223c71464e409116… |
| [Anthropic] 2025 Type 2 SOC 3 Report.pdf | 4a2eec78000029cc… |
| [Anthropic] Anthropic's Enterprise Security Posture.pdf | ae3a40f6cd75a651… |
| [Anthropic] Data Handling & Key Management.pdf | 4d86c308354881f8… |
| [Anthropic] Data Loss Prevention & Content Controls.pdf | cd85d044e4ccf41f… |
| [Anthropic] ISO 27001 Certificate (2025).pdf | b8d621cc3c9ac880… |
| [Anthropic] ISO 42001 Certificate (2025).pdf | a72bbe24b44b5c17… |
| [Anthropic] Identity & Access Controls.pdf | 1b9f0c3575e892b5… |
| [Anthropic] Isolation & Connectivity.pdf | 092b7815b2e4ca76… |
| c4ecertificationpackageoverview.clean.json | 2feb224ff02db96d… |
| v1.0 Claude Code FISMA Best Practices.pdf | 1c4ab4052647fe75… |
===== Claude Mythos 5 and Claude Fable 5 Training Data Summary .pdf ===== Public Summary of Training Content Claude Mythos 5 & Claude Fable 5 Training Data Summary Version of the Summary: Version #1 Last update: July 24, 2026 1. General information 1.1. Provider identification Provider name and contact details: Anthropic Ireland, Limited 6th Floor South Bank House, Barrow Street, Dublin 4, Dublin Ireland Authorised representative name and contact details: Not applicable 1.2. Model identification Versioned model name(s): Claude Mythos 5 and Claude Fable 5 Model Card: www.anthropic.com/system-cards Model dependencies: Not applicable Date of placement of the model on the Union market: June 9, 2026 1.3 Modalities, overall training data size and other characteristics Modality Select the modalities present in the training data, to the extent that they are identifiable Training data size For each selected modality, select the range within which the estimated total training data size for that modality falls. Dynamic datasets may be excluded from the estimation. Types of content For each selected modality, provide a general description of the type of content that has been included in the training data. Text ☐ Less than 1 billion tokens ☐ 1billion to 10 trillions tokens X More than 10 trillions tokens The training corpus for the model includes an array of text types, including short and long-form texts, software code, synthetic text, prose in a variety of languages, mathematical data, and prompts and preference data used during reinforcement learning. Image ☐ Less than 1 million images ☐ 1Million to1 billion images X More than 1 billion images The training corpus for the model includes an array of image types, including photographs, interleaved 2 Training Data Summary text and images from websites, computer graphics, and prompts and preference data used during reinforcement learning. Video Not applicable Audio Not applicable Other Not applicable Latest date of data acquisition/collection for model training: A number of different datasets, with varying publication and cut-off dates, are included in the training corpus, with some data being acquired/collected up to April 2026. Description of the linguistic characteristics of the overall training data: Training sources deliberately include a diverse range of global languages, both European and non-European, including those with relatively high numbers of speakers ( e.g., English, Chinese, French, Spanish) as well as comparably lower volumes ( e.g., Basque, Breton, Korean). 2. List of data sources 2. List of data sources 2.1. Publicly available datasets Have you used publicly available datasets to train the model? Yes If yes, specify the modality(ies) of the content covered by the datasets concerned: Text, Image List of large publicly available datasets: The training corpus is derived from several publicly accessible repositories, notably including Common Crawl, a repository of web crawl data, as well as specialized datasets available through platforms like GitHub and HuggingFace. General description of other publicly available datasets not listed above: Data from other publicly available datasets is included in the training corpus, including mathematical data and some image and caption datasets. 2.2 Private non-publicly available datasets obtained from third parties 2.2.1. Datasets commercially licensed by rightsholders or their representatives Have you concluded transactional commercial licensing agreement(s) with rightsholder(s) or with their representatives? Yes 3 Training Data Summary If yes, specify the modality(ies) of the content covered by the datasets concerned: Text, Image 2.2.2. Private datasets obtained from other third parties Have you obtained private datasets from third parties that are not licensed as described in Section 2.2.1, such as data obtained from providers of private databases, or data intermediaries? Yes If yes, specify the modality(ies) of the content covered by the datasets concerned: Text, Image If publicly known, list private datasets obtained from other third parties: Not applicable General description of non-publicly known private datasets obtained from third parties We obtain non-publicly known private datasets from third parties covering diverse domains and content types. 2.3 Data crawled and scraped from online sources Were crawlers used by the provider or on behalf of? Yes If yes, specify crawler name(s)/identifier(s): ClaudeBot Purposes of the crawler(s): ClaudeBot collects web content that could potentially contribute to the model’s training. General description of crawler behaviour: We aim to minimize disruption to website owners and be thoughtful about how quickly ClaudeBot crawls domains, including by respecting crawl-delay and disallow directives in robots.txt files, where appropriate, and respecting anti-circumvention technologies such as paywalls, password protection, and CAPTCHAs. More information on our crawlers and how they work can be found at: https://support.claude.com/en/articles/8896518-does-anthropic-crawl-data-from-the-web-and-how-can-site-owners-block-the-crawler. Period of data collection: March 2024 - April 2026 Comprehensive description of the type of content and online sources crawled: The crawlers may be exposed to a wide variety of content and online sources, including most forms of publicly available online data. 4 Training Data Summary Type of modality covered: Text, Image Summary of the most relevant domain names crawled: The portion of the model's training corpus derived from data crawled and scraped from online sources includes technical documentation, open-source software, predominantly text-based reference sites, document sharing sites, and math sites. Top-level domains such as .com, .org, and .net are included alongside sites from a range of different countries. 2.4 User data Was data from user interactions with the AI model (e.g. user input and prompts) used to train the model? Yes Was data collected from user interactions with the provider’s other services or products used to train the model? Yes If yes, provide a general description of the provider’s services or products that were used to collect the user data: To the extent permitted by Anthropic’s terms of service, privacy policy, and other contracts, and in line with applicable law, if a user explicitly reports feedback or bugs to us (e.g., via thumbs and feedback buttons) or otherwise chooses to allow us to use their data, then chats and coding session data may be used in model training. More information is available at https://privacy.claude.com/en/articles/7996868-is-my-data-used-for-model-training We may also incorporate data derived from Anthropic employees’ use of internal-only model versions. Type of modality covered: Text 2.5 Synthetic data Was synthetic AI-generated data created by the provider or on their behalf to train the model? Yes If yes, modality of the synthetic data: Text, Image 5 Training Data Summary If yes, specify the general-purpose AI model(s) used to generate the synthetic data if available on the market: Synthetic data was provided by speech to text models, large language models (LLM), and vision-language models (VLM). Information about other AI models, including provider’s own AI model(s) not available on the market, used to generate synthetic data to train the model to which this Summary applies: Some synthetic data used in training was generated by Anthropic models not available on the market. 2.6 Other sources of data Have data sources other than those described in Sections 2.1 to 2.5 been used to train the model? Yes If yes, provide a narrative description of these data sources and the data: A portion of the data corpus comes from acquired physical texts. 1. Data processing aspects 3. Data processing aspects 3.1. Respect of reservation of rights from text and data mining exception or limitation Are you a Signatory to the Code of Practice for general-purpose AI models that includes commitments to respect reservations of rights from the TDM exception or limitation? Yes Describe the measures implemented before model training to respect reservations of rights from the TDM exception or limitation before and during data collection, including the opt-out protocols and solutions honoured by the provider or, as applicable, by third parties from which datasets have been obtained: ClaudeBot respects crawl-delay and disallow directives in robots.txt files where appropriate, as well as anti-circumvention technologies such as paywalls, password protection, and CAPTCHAs. More information on our crawlers and how they work can be 6 Training Data Summary found at: https://support.claude.com/en/articles/8896518-does-anthropic-crawl-data-from-the-web-and-how-can-site-owners-block-the-crawler. 3.2 Removal of illegal content General description of measures taken: We take a number of protective measures to remove illegal content from the training corpus such as active filtering, scoring, moderation, and blocking. 3.3. Other information (optional) Other relevant information about data processing (optional): Not applicable 7 ===== Claude Opus 4.7 Training Data Summary .pdf ===== Public Summary of Training Content Claude Opus 4.7 Training Data Summary Version of the Summary: Version #1 Last update: July 24, 2026 General information 1. General information 1.1. Provider identification Provider name and contact details: Anthropic Ireland, Limited 6th Floor South Bank House, Barrow Street, Dublin 4, Dublin Ireland Authorised representative name and contact details: Not applicable 1.2. Model identification Versioned model name(s): Claude Opus 4.7 Model Card: www.anthropic.com/system-cards Model dependencies: Not applicable Date of placement of the model on the Union market: April 16, 2026 1.3 Modalities, overall training data size and other characteristics Modality Select the modalities present in the training data, to the extent that they are identifiable Training data size For each selected modality, select the range within which the estimated total training data size for that modality falls. Dynamic datasets may be excluded from the estimation. Types of content For each selected modality, provide a general description of the type of content that has been included in the training data. Text ☐ Less than 1 billion tokens ☐ 1billion to 10 trillions tokens X More than 10 trillions tokens The training corpus for the model includes an array of text types, including short and long-form texts, software code, synthetic text, prose in a variety of languages, mathematical data, and prompts and preference data used during reinforcement learning. 2 Training Data Summary Image ☐ Less than 1 million images ☐ 1Million to1 billion images X More than 1 billion images The training corpus for the model includes an array of image types, including photographs, interleaved text and images from websites, computer graphics, and prompts and preference data used during reinforcement learning. Video Not applicable Audio Not applicable Other Not applicable Latest date of data acquisition/collection for model training: A number of different datasets, with varying publication and cut-off dates, are included in the training corpus, with some data being acquired/collected up to April 2026. Description of the linguistic characteristics of the overall training data: Training sources deliberately include a diverse range of global languages, both European and non-European, including those with relatively high numbers of speakers ( e.g., English, Chinese, French, Spanish) as well as comparably lower volumes ( e.g., Basque, Breton, Korean). 2. List of data sources 2. List of data sources 2.1. Publicly available datasets Have you used publicly available datasets to train the model? Yes If yes, specify the modality(ies) of the content covered by the datasets concerned: Text, Image List of large publicly available datasets: The training corpus is derived from several publicly accessible repositories, notably including Common Crawl, a repository of web crawl data, as well as specialized datasets available through platforms like GitHub and HuggingFace. General description of other publicly available datasets not listed above: Data from other publicly available datasets is included in the training corpus, including mathematical data and some image and caption datasets. 2.2 Private non-publicly available datasets obtained from third parties 3 Training Data Summary 2.2.1. Datasets commercially licensed by rightsholders or their representatives Have you concluded transactional commercial licensing agreement(s) with rightsholder(s) or with their representatives? Yes If yes, specify the modality(ies) of the content covered by the datasets concerned: Text, Image 2.2.2. Private datasets obtained from other third parties Have you obtained private datasets from third parties that are not licensed as described in Section 2.2.1, such as data obtained from providers of private databases, or data intermediaries? Yes If yes, specify the modality(ies) of the content covered by the datasets concerned: Text, Image If publicly known, list private datasets obtained from other third parties: Not applicable General description of non-publicly known private datasets obtained from third parties We obtain non-publicly known private datasets from third parties covering diverse domains and content types. 2.3 Data crawled and scraped from online sources Were crawlers used by the provider or on behalf of? Yes If yes, specify crawler name(s)/identifier(s): ClaudeBot Purposes of the crawler(s): ClaudeBot collects web content that could potentially contribute to the model’s training. General description of crawler behaviour: We aim to minimize disruption to website owners and be thoughtful about how quickly ClaudeBot crawls domains, including by respecting crawl-delay and disallow directives in robots.txt files, where appropriate, and respecting anti-circumvention technologies such as paywalls, password protection, and CAPTCHAs. More information on our crawlers and how they work can be found at: https://support.claude.com/en/articles/8896518-does-anthropic-crawl-data-from-the-web-and-how-can-site-owners-block-the-crawler. 4 Training Data Summary Period of data collection: March 2024 - April 2026 Comprehensive description of the type of content and online sources crawled: The crawlers may be exposed to a wide variety of content and online sources, including most forms of publicly available online data. Type of modality covered: Text, Image Summary of the most relevant domain names crawled: The portion of the model's training corpus derived from data crawled and scraped from online sources includes technical documentation, open-source software, predominantly text-based reference sites, document sharing sites, and math sites. Top-level domains such as .com, .org, and .net are included alongside sites from a range of different countries. 2.4 User data Was data from user interactions with the AI model (e.g. user input and prompts) used to train the model? Yes Was data collected from user interactions with the provider’s other services or products used to train the model? Yes If yes, provide a general description of the provider’s services or products that were used to collect the user data: To the extent permitted by Anthropic’s terms of service, privacy policy, and other contracts, and in line with applicable law, if a user explicitly reports feedback or bugs to us (e.g., via thumbs and feedback buttons) or otherwise chooses to allow us to use their data, then chats and coding session data may be used in model training. More information is available at https://privacy.claude.com/en/articles/7996868-is-my-data-used-for-model-training We may also incorporate data derived from Anthropic employees’ use of internal-only model versions. Type of modality covered: Text 2.5 Synthetic data Was synthetic AI-generated data created by the Yes 5 Training Data Summary provider or on their behalf to train the model? If yes, modality of the synthetic data: Text, Image If yes, specify the general-purpose AI model(s) used to generate the synthetic data if available on the market: Synthetic data was provided by speech to text models, large language models (LLM), and vision-language models (VLM). Information about other AI models, including provider’s own AI model(s) not available on the market, used to generate synthetic data to train the model to which this Summary applies: Some synthetic data used in training was generated by Anthropic models not available on the market. 2.6 Other sources of data Have data sources other than those described in Sections 2.1 to 2.5 been used to train the model? Yes If yes, provide a narrative description of these data sources and the data: A portion of the data corpus comes from acquired physical texts. 1. Data processing aspects 3. Data processing aspects 3.1. Respect of reservation of rights from text and data mining exception or limitation Are you a Signatory to the Code of Practice for general-purpose AI models that includes commitments to respect reservations of rights from the TDM exception or limitation? Yes 6 Training Data Summary Describe the measures implemented before model training to respect reservations of rights from the TDM exception or limitation before and during data collection, including the opt-out protocols and solutions honoured by the provider or, as applicable, by third parties from which datasets have been obtained: ClaudeBot respects crawl-delay and disallow directives in robots.txt files where appropriate, as well as anti-circumvention technologies such as paywalls, password protection, and CAPTCHAs. More information on our crawlers and how they work can be found at: https://support.claude.com/en/articles/8896518-does-anthropic-crawl-data-from-the-web-and-how-can-site-owners-block-the-crawler. 3.2 Removal of illegal content General description of measures taken: We take a number of protective measures to remove illegal content from the training corpus such as active filtering, scoring, moderation, and blocking. 3.3. Other information (optional) Other relevant information about data processing (optional): Not applicable 7 ===== Claude Opus 4.8 Training Data Summary .pdf ===== Public Summary of Training Content Claude Opus 4.8 Training Data Summary Version of the Summary: Version #1 Last update: July 24, 2026 General information 1. General information 1.1. Provider identification Provider name and contact details: Anthropic Ireland, Limited 6th Floor South Bank House, Barrow Street, Dublin 4, Dublin Ireland Authorised representative name and contact details: Not applicable 1.2. Model identification Versioned model name(s): Claude Opus 4.8 Model Card: www.anthropic.com/system-cards Model dependencies: Not applicable Date of placement of the model on the Union market: May 28, 2026 1.3 Modalities, overall training data size and other characteristics Modality Select the modalities present in the training data, to the extent that they are identifiable Training data size For each selected modality, select the range within which the estimated total training data size for that modality falls. Dynamic datasets may be excluded from the estimation. Types of content For each selected modality, provide a general description of the type of content that has been included in the training data. Text ☐ Less than 1 billion tokens ☐ 1billion to 10 trillions tokens X More than 10 trillions tokens The training corpus for the model includes an array of text types, including short and long-form texts, software code, synthetic text, prose in a variety of languages, mathematical data, and prompts and preference data used during reinforcement learning. Image ☐ Less than 1 million images ☐ 1Million to1 billion images The training corpus for the model includes an array of image types, 2 Training Data Summary X More than 1 billion images including photographs, interleaved text and images from websites, computer graphics, and prompts and preference data used during reinforcement learning. Video Not applicable Audio Not applicable Other Not applicable Latest date of data acquisition/collection for model training: A number of different datasets, with varying publication and cut-off dates, are included in the training corpus, with some data being acquired/collected up to May 2026. Description of the linguistic characteristics of the overall training data: Training sources deliberately include a diverse range of global languages, both European and non-European, including those with relatively high numbers of speakers ( e.g., English, Chinese, French, Spanish) as well as comparably lower volumes ( e.g., Basque, Breton, Korean). 2. List of data sources 2. List of data sources 2.1. Publicly available datasets Have you used publicly available datasets to train the model? Yes If yes, specify the modality(ies) of the content covered by the datasets concerned: Text, Image List of large publicly available datasets: The training corpus is derived from several publicly accessible repositories, notably including Common Crawl, a repository of web crawl data, as well as specialized datasets available through platforms like GitHub and HuggingFace. General description of other publicly available datasets not listed above: Data from other publicly available datasets is included in the training corpus, including mathematical data and some image and caption datasets. 2.2 Private non-publicly available datasets obtained from third parties 2.2.1. Datasets commercially licensed by rightsholders or their representatives Have you concluded transactional commercial licensing agreement(s) with rightsholder(s) or with their representatives? Yes 3 Training Data Summary If yes, specify the modality(ies) of the content covered by the datasets concerned: Text, Image 2.2.2. Private datasets obtained from other third parties Have you obtained private datasets from third parties that are not licensed as described in Section 2.2.1, such as data obtained from providers of private databases, or data intermediaries? Yes If yes, specify the modality(ies) of the content covered by the datasets concerned: Text, Image If publicly known, list private datasets obtained from other third parties: Not applicable General description of non-publicly known private datasets obtained from third parties We obtain non-publicly known private datasets from third parties covering diverse domains and content types. 2.3 Data crawled and scraped from online sources Were crawlers used by the provider or on behalf of? Yes If yes, specify crawler name(s)/identifier(s): ClaudeBot Purposes of the crawler(s): ClaudeBot collects web content that could potentially contribute to the model’s training. General description of crawler behaviour: We aim to minimize disruption to website owners and be thoughtful about how quickly ClaudeBot crawls domains, including by respecting crawl-delay and disallow directives in robots.txt files, where appropriate, and respecting anti-circumvention technologies such as paywalls, password protection, and CAPTCHAs. More information on our crawlers and how they work can be found at: https://support.claude.com/en/articles/8896518-does-anthropic-crawl-data-from-the-web-and-how-can-site-owners-block-the-crawler. Period of data collection: March 2024 - May 2026 Comprehensive description of the type of content and online sources crawled: The crawlers may be exposed to a wide variety of content and online sources, including most forms of publicly available online data. 4 Training Data Summary Type of modality covered: Text, Image Summary of the most relevant domain names crawled: The portion of the model's training corpus derived from data crawled and scraped from online sources includes technical documentation, open-source software, predominantly text-based reference sites, document sharing sites, and math sites. Top-level domains such as .com, .org, and .net are included alongside sites from a range of different countries. 2.4 User data Was data from user interactions with the AI model (e.g. user input and prompts) used to train the model? Yes Was data collected from user interactions with the provider’s other services or products used to train the model? Yes If yes, provide a general description of the provider’s services or products that were used to collect the user data: To the extent permitted by Anthropic’s terms of service, privacy policy, and other contracts, and in line with applicable law, if a user explicitly reports feedback or bugs to us (e.g., via thumbs and feedback buttons) or otherwise chooses to allow us to use their data, then chats and coding session data may be used in model training. More information is available at https://privacy.claude.com/en/articles/7996868-is-my-data-used-for-model-training We may also incorporate data derived from Anthropic employees’ use of internal-only model versions. Type of modality covered: Text 2.5 Synthetic data Was synthetic AI-generated data created by the provider or on their behalf to train the model? Yes If yes, modality of the synthetic data: Text, Image 5 Training Data Summary If yes, specify the general-purpose AI model(s) used to generate the synthetic data if available on the market: Synthetic data was provided by speech to text models, large language models (LLM), and vision-language models (VLM). Information about other AI models, including provider’s own AI model(s) not available on the market, used to generate synthetic data to train the model to which this Summary applies: Some synthetic data used in training was generated by Anthropic models not available on the market. 2.6 Other sources of data Have data sources other than those described in Sections 2.1 to 2.5 been used to train the model? Yes If yes, provide a narrative description of these data sources and the data: A portion of the data corpus comes from acquired physical texts. 1. Data processing aspects 3. Data processing aspects 3.1. Respect of reservation of rights from text and data mining exception or limitation Are you a Signatory to the Code of Practice for general-purpose AI models that includes commitments to respect reservations of rights from the TDM exception or limitation? Yes Describe the measures implemented before model training to respect reservations of rights from the TDM exception or limitation before and during data collection, including the opt-out protocols and solutions honoured by the provider or, as applicable, by third parties from which datasets have been obtained: ClaudeBot respects crawl-delay and disallow directives in robots.txt files where appropriate, as well as anti-circumvention technologies such as paywalls, password protection, and CAPTCHAs. More information on our crawlers and how they work can be 6 Training Data Summary found at: https://support.claude.com/en/articles/8896518-does-anthropic-crawl-data-from-the-web-and-how-can-site-owners-block-the-crawler. 3.2 Removal of illegal content General description of measures taken: We take a number of protective measures to remove illegal content from the training corpus such as active filtering, scoring, moderation, and blocking. 3.3. Other information (optional) Other relevant information about data processing (optional): Not applicable 7 ===== Claude Opus 5 Training Data Summary .pdf ===== Public Summary of Training Content Claude Opus 5 Training Data Summary Version of the Summary: Version #1 Last update: July 23, 2026 General information 1. General information 1.1. Provider identification Provider name and contact details: Anthropic Ireland, Limited 6th Floor South Bank House, Barrow Street, Dublin 4, Dublin Ireland Authorised representative name and contact details: Not applicable 1.2. Model identification Versioned model name(s): Claude Opus 5 Model Card: www.anthropic.com/system-cards Model dependencies: Not applicable Date of placement of the model on the Union market: July 23, 2026 1.3 Modalities, overall training data size and other characteristics Modality Select the modalities present in the training data, to the extent that they are identifiable Training data size For each selected modality, select the range within which the estimated total training data size for that modality falls. Dynamic datasets may be excluded from the estimation. Types of content For each selected modality, provide a general description of the type of content that has been included in the training data. 2 Training Data Summary Text ☐ Less than 1 billion tokens ☐ 1billion to 10 trillions tokens X More than 10 trillions tokens The training corpus for the model includes an array of text types, including short and long-form texts, software code, synthetic text, prose in a variety of languages, mathematical data, and prompts and preference data used during reinforcement learning. Image ☐ Less than 1 million images ☐ 1Million to1 billion images X More than 1 billion images The training corpus for the model includes an array of image types, including photographs, interleaved text and images from websites, computer graphics, and prompts and preference data used during reinforcement learning. Video Not applicable Audio Not applicable Other Not applicable Latest date of data acquisition/collection for model training: A number of different datasets, with varying publication and cut-off dates, are included in the training corpus, with some data being acquired/collected up to July 2026. Description of the linguistic characteristics of the overall training data: Training sources deliberately include a diverse range of global languages, both European and non-European, including those with relatively high numbers of speakers ( e.g., English, Chinese, French, Spanish) as well as comparably lower volumes ( e.g., Basque, Breton, Korean). 2. List of data sources 2. List of data sources 2.1. Publicly available datasets Have you used publicly available datasets to train the model? Yes If yes, specify the modality(ies) of the content covered by the datasets concerned: Text, Image List of large publicly available datasets: The training corpus is derived from several publicly accessible repositories, notably including Common Crawl, a 3 Training Data Summary repository of web crawl data, as well as specialized datasets available through platforms like GitHub and HuggingFace. General description of other publicly available datasets not listed above: Data from other publicly available datasets is included in the training corpus, including mathematical data and some image and caption datasets. 2.2 Private non-publicly available datasets obtained from third parties 2.2.1. Datasets commercially licensed by rightsholders or their representatives Have you concluded transactional commercial licensing agreement(s) with rightsholder(s) or with their representatives? Yes If yes, specify the modality(ies) of the content covered by the datasets concerned: Text, Image 2.2.2. Private datasets obtained from other third parties Have you obtained private datasets from third parties that are not licensed as described in Section 2.2.1, such as data obtained from providers of private databases, or data intermediaries? Yes If yes, specify the modality(ies) of the content covered by the datasets concerned: Text, Image If publicly known, list private datasets obtained from other third parties: Not applicable General description of non-publicly known private datasets obtained from third parties We obtain non-publicly known private datasets from third parties covering diverse domains and content types. 2.3 Data crawled and scraped from online sources Were crawlers used by the provider or on behalf of? Yes 4 Training Data Summary If yes, specify crawler name(s)/identifier(s): ClaudeBot Purposes of the crawler(s): ClaudeBot collects web content that could potentially contribute to the model’s training. General description of crawler behaviour: We aim to minimize disruption to website owners and be thoughtful about how quickly ClaudeBot crawls domains, including by respecting crawl-delay and disallow directives in robots.txt files, where appropriate, and respecting anti-circumvention technologies such as paywalls, password protection, and CAPTCHAs. More information on our crawlers and how they work can be found at: https://support.claude.com/en/articles/8896518-does-anthropic-crawl-data-from-the-web-and-how-can-site-owners-block-the-crawler. Period of data collection: March 2024 - July 2026 Comprehensive description of the type of content and online sources crawled: The crawlers may be exposed to a wide variety of content and online sources, including most forms of publicly available online data. Type of modality covered: Text, Image Summary of the most relevant domain names crawled: The portion of the model's training corpus derived from data crawled and scraped from online sources includes technical documentation, open-source software, predominantly text-based reference sites, document sharing sites, and math sites. Top-level domains such as .com, .org, and .net are included alongside sites from a range of different countries. 2.4 User data Was data from user interactions with the AI model (e.g. user input and prompts) used to train the model? Yes Was data collected from user interactions with the provider’s other services or products used to train the model? Yes If yes, provide a general description of the provider’s services or products that were used to collect the user data: To the extent permitted by Anthropic’s terms of service, privacy policy, and other contracts, and in line with applicable law, if a 5 Training Data Summary user explicitly reports feedback or bugs to us (e.g., via thumbs and feedback buttons) or otherwise chooses to allow us to use their data, then chats and coding session data may be used in model training. More information is available at https://privacy.claude.com/en/articles/7996868-is-my-data-used-for-model-training We may also incorporate data derived from Anthropic employees’ use of internal-only model versions. Type of modality covered: Text 2.5 Synthetic data Was synthetic AI-generated data created by the provider or on their behalf to train the model? Yes If yes, modality of the synthetic data: Text, Image If yes, specify the general-purpose AI model(s) used to generate the synthetic data if available on the market: Synthetic data was provided by speech to text models, large language models (LLM), and vision-language models (VLM). Information about other AI models, including provider’s own AI model(s) not available on the market, used to generate synthetic data to train the model to which this Summary applies: Some synthetic data used in training was generated by Anthropic models not available on the market. 2.6 Other sources of data Have data sources other than those described in Sections 2.1 to 2.5 been used to train the model? Yes 6 Training Data Summary If yes, provide a narrative description of these data sources and the data: A portion of the data corpus comes from acquired physical texts. 1. Data processing aspects 3. Data processing aspects 3.1. Respect of reservation of rights from text and data mining exception or limitation Are you a Signatory to the Code of Practice for general-purpose AI models that includes commitments to respect reservations of rights from the TDM exception or limitation? Yes Describe the measures implemented before model training to respect reservations of rights from the TDM exception or limitation before and during data collection, including the opt-out protocols and solutions honoured by the provider or, as applicable, by third parties from which datasets have been obtained: ClaudeBot respects crawl-delay and disallow directives in robots.txt files where appropriate, as well as anti-circumvention technologies such as paywalls, password protection, and CAPTCHAs. More information on our crawlers and how they work can be found at: https://support.claude.com/en/articles/8896518-does-anthropic-crawl-data-from-the-web-and-how-can-site-owners-block-the-crawler. 3.2 Removal of illegal content General description of measures taken: We take a number of protective measures to remove illegal content from the training corpus such as active filtering, scoring, moderation, and blocking. 3.3. Other information (optional) Other relevant information about data processing (optional): Not applicable 7 Training Data Summary 8 ===== Claude Sonnet 5 Training Data Summary .pdf ===== Public Summary of Training Content Claude Sonnet 5 Training Data Summary Version of the Summary: Version #1 Last update: July 24, 2026 General information\ 1. General information 1.1. Provider identification Provider name and contact details: Anthropic Ireland, Limited 6th Floor South Bank House, Barrow Street, Dublin 4, Dublin Ireland Authorised representative name and contact details: Not applicable 1.2. Model identification Versioned model name(s): Claude Sonnet 5 Model Card: www.anthropic.com/system-cards Model dependencies: Not applicable Date of placement of the model on the Union market: June 30, 2026 1.3 Modalities, overall training data size and other characteristics Modality Select the modalities present in the training data, to the extent that they are identifiable Training data size For each selected modality, select the range within which the estimated total training data size for that modality falls. Dynamic datasets may be excluded from the estimation. Types of content For each selected modality, provide a general description of the type of content that has been included in the training data. Text ☐ Less than 1 billion tokens ☐ 1billion to 10 trillions tokens X More than 10 trillions tokens The training corpus for the model includes an array of text types, including short and long-form texts, software code, synthetic text, prose in a variety of languages, mathematical data, and prompts and preference data used during reinforcement learning. 2 Training Data Summary Image ☐ Less than 1 million images ☐ 1Million to1 billion images X More than 1 billion images The training corpus for the model includes an array of image types, including photographs, interleaved text and images from websites, computer graphics, and prompts and preference data used during reinforcement learning. Video Not applicable Audio Not applicable Other Not applicable Latest date of data acquisition/collection for model training: A number of different datasets, with varying publication and cut-off dates, are included in the training corpus, with some data being acquired/collected up to May 2026. Description of the linguistic characteristics of the overall training data: Training sources deliberately include a diverse range of global languages, both European and non-European, including those with relatively high numbers of speakers ( e.g., English, Chinese, French, Spanish) as well as comparably lower volumes ( e.g., Basque, Breton, Korean). 2. List of data sources 2. List of data sources 2.1. Publicly available datasets Have you used publicly available datasets to train the model? Yes If yes, specify the modality(ies) of the content covered by the datasets concerned: Text, Image List of large publicly available datasets: The training corpus is derived from several publicly accessible repositories, notably including Common Crawl, a repository of web crawl data, as well as specialized datasets available through platforms like GitHub and HuggingFace. General description of other publicly available datasets not listed above: Data from other publicly available datasets is included in the training corpus, including mathematical data and some image and caption datasets. 2.2 Private non-publicly available datasets obtained from third parties 3 Training Data Summary 2.2.1. Datasets commercially licensed by rightsholders or their representatives Have you concluded transactional commercial licensing agreement(s) with rightsholder(s) or with their representatives? Yes If yes, specify the modality(ies) of the content covered by the datasets concerned: Text, Image 2.2.2. Private datasets obtained from other third parties Have you obtained private datasets from third parties that are not licensed as described in Section 2.2.1, such as data obtained from providers of private databases, or data intermediaries? Yes If yes, specify the modality(ies) of the content covered by the datasets concerned: Text, Image If publicly known, list private datasets obtained from other third parties: Not applicable General description of non-publicly known private datasets obtained from third parties We obtain non-publicly known private datasets from third parties covering diverse domains and content types. 2.3 Data crawled and scraped from online sources Were crawlers used by the provider or on behalf of? Yes If yes, specify crawler name(s)/identifier(s): ClaudeBot Purposes of the crawler(s): ClaudeBot collects web content that could potentially contribute to the model’s training. General description of crawler behaviour: We aim to minimize disruption to website owners and be thoughtful about how quickly ClaudeBot crawls domains, including by respecting crawl-delay and disallow directives in robots.txt files, where appropriate, and respecting anti-circumvention technologies such as paywalls, password protection, and CAPTCHAs. More information on our crawlers and how they work can be found at: https://support.claude.com/en/articles/8896518-does-anthropic-crawl-data-from-the-web-and-how-can-site-owners-block-the-crawler. 4 Training Data Summary Period of data collection: March 2024 - May 2026 Comprehensive description of the type of content and online sources crawled: The crawlers may be exposed to a wide variety of content and online sources, including most forms of publicly available online data. Type of modality covered: Text, Image Summary of the most relevant domain names crawled: The portion of the model's training corpus derived from data crawled and scraped from online sources includes technical documentation, open-source software, predominantly text-based reference sites, document sharing sites, and math sites. Top-level domains such as .com, .org, and .net are included alongside sites from a range of different countries. 2.4 User data Was data from user interactions with the AI model (e.g. user input and prompts) used to train the model? Yes Was data collected from user interactions with the provider’s other services or products used to train the model? Yes If yes, provide a general description of the provider’s services or products that were used to collect the user data: To the extent permitted by Anthropic’s terms of service, privacy policy, and other contracts, and in line with applicable law, if a user explicitly reports feedback or bugs to us (e.g., via thumbs and feedback buttons) or otherwise chooses to allow us to use their data, then chats and coding session data may be used in model training. More information is available at https://privacy.claude.com/en/articles/7996868-is-my-data-used-for-model-training We may also incorporate data derived from Anthropic employees’ use of internal-only model versions. Type of modality covered: Text 2.5 Synthetic data Was synthetic AI-generated data created by the Yes 5 Training Data Summary provider or on their behalf to train the model? If yes, modality of the synthetic data: Text, Image If yes, specify the general-purpose AI model(s) used to generate the synthetic data if available on the market: Synthetic data was provided by speech to text models, large language models (LLM), and vision-language models (VLM). Information about other AI models, including provider’s own AI model(s) not available on the market, used to generate synthetic data to train the model to which this Summary applies: Some synthetic data used in training was generated by Anthropic models not available on the market. 2.6 Other sources of data Have data sources other than those described in Sections 2.1 to 2.5 been used to train the model? Yes If yes, provide a narrative description of these data sources and the data: A portion of the data corpus comes from acquired physical texts. 1. Data processing aspects 3. Data processing aspects 3.1. Respect of reservation of rights from text and data mining exception or limitation Are you a Signatory to the Code of Practice for general-purpose AI models that includes commitments to respect reservations of rights from the TDM exception or limitation? Yes 6 Training Data Summary Describe the measures implemented before model training to respect reservations of rights from the TDM exception or limitation before and during data collection, including the opt-out protocols and solutions honoured by the provider or, as applicable, by third parties from which datasets have been obtained: ClaudeBot respects crawl-delay and disallow directives in robots.txt files where appropriate, as well as anti-circumvention technologies such as paywalls, password protection, and CAPTCHAs. More information on our crawlers and how they work can be found at: https://support.claude.com/en/articles/8896518-does-anthropic-crawl-data-from-the-web-and-how-can-site-owners-block-the-crawler. 3.2 Removal of illegal content General description of measures taken: We take a number of protective measures to remove illegal content from the training corpus such as active filtering, scoring, moderation, and blocking. 3.3. Other information (optional) Other relevant information about data processing (optional): Not applicable 7 ===== Claude Training Data Summary (Fable 5.1 + Mythos 5.1) PDF.pdf ===== Public Summary of Training Content Claude Fable 5.1 & Claude Mythos 5.1 Training Data Summary Version of the Summary: Version #1 Last update: September 1, 2026 1. General information 1.1. Provider identification Provider name and contact details: Anthropic Ireland, Limited 6th Floor South Bank House, Barrow Street, Dublin 4, Dublin Ireland Authorised representative name and contact details: Not applicable 1.2. Model identification Versioned model name(s): Claude Fable 5.1 Claude Mythos 5.1 Claude Fable 5.1 and Claude Mythos 5.1 are both versions of the same underlying model, with differing safeguard configurations and access: Claude Fable 5.1 is the generally available version with full safeguards, whereas Claude Mythos 5.1 has reduced safeguards and is available only through trusted access programs. Model Card: www.anthropic.com/system-cards Model dependencies: Not applicable Date of placement of the model on the Union market: September 1, 2026 1.3 Modalities, overall training data size and other characteristics Modality Select the modalities present in the training data, to the extent that Training data size For each selected modality, select the range within which the estimated total training data size for that modality falls. Dynamic datasets may be excluded from the estimation. Types of content For each selected modality, provide a general description of the type of content that has been included in the training data. 2 Training Data Summary they are identifiable Text ☐ Less than 1 billion tokens ☐ 1billion to 10 trillions tokens X More than 10 trillions tokens The training corpus for the model includes an array of text types, including short and long-form texts, software code, synthetic text, prose in a variety of languages, mathematical data, and prompts and preference data used during reinforcement learning. Image ☐ Less than 1 million images ☐ 1Million to1 billion images X More than 1 billion images The training corpus for the model includes an array of image types, including photographs, interleaved text and images from websites, computer graphics, and prompts and preference data used during reinforcement learning. Video Not applicable Audio Not applicable Other Not applicable Latest date of data acquisition/collection for model training: A number of different datasets, with varying publication and cut-off dates, are included in the training corpus, with some data being acquired/collected up to August 2026. Description of the linguistic characteristics of the overall training data: Training sources deliberately include a diverse range of global languages, both European and non-European, including those with relatively high numbers of speakers ( e.g., English, Chinese, French, Spanish) as well as comparably lower volumes ( e.g., Basque, Breton, Korean). 2. List of data sources 2. List of data sources 2.1. Publicly available datasets Have you used publicly available datasets to train the model? Yes If yes, specify the modality(ies) of the content covered by the datasets concerned: Text, Image List of large publicly available datasets: The training corpus is derived from several publicly accessible repositories, notably including Common Crawl, a repository of web crawl data, as well as specialized datasets available through platforms like GitHub and HuggingFace. 3 Training Data Summary General description of other publicly available datasets not listed above: Data from other publicly available datasets is included in the training corpus, including mathematical data and some image and caption datasets. 2.2 Private non-publicly available datasets obtained from third parties 2.2.1. Datasets commercially licensed by rightsholders or their representatives Have you concluded transactional commercial licensing agreement(s) with rightsholder(s) or with their representatives? Yes If yes, specify the modality(ies) of the content covered by the datasets concerned: Text, Image 2.2.2. Private datasets obtained from other third parties Have you obtained private datasets from third parties that are not licensed as described in Section 2.2.1, such as data obtained from providers of private databases, or data intermediaries? Yes If yes, specify the modality(ies) of the content covered by the datasets concerned: Text, Image If publicly known, list private datasets obtained from other third parties: Not applicable General description of non-publicly known private datasets obtained from third parties We obtain non-publicly known private datasets from third parties covering diverse domains and content types. 2.3 Data crawled and scraped from online sources Were crawlers used by the provider or on behalf of? Yes If yes, specify crawler name(s)/identifier(s): ClaudeBot Purposes of the crawler(s): ClaudeBot collects web content that could potentially contribute to the model’s training. General description of crawler behaviour: We aim to minimize disruption to website owners and be thoughtful about how quickly ClaudeBot crawls domains, including by respecting crawl-delay and 4 Training Data Summary disallow directives in robots.txt files, where appropriate, and respecting anti-circumvention technologies such as paywalls, password protection, and CAPTCHAs. More information on our crawlers and how they work can be found at: https://support.claude.com/en/articles/8896518-does-anthropic-crawl-data-from-the-web-and-how-can-site-owners-block-the-crawler. Period of data collection: March 2024 - July 2026 Comprehensive description of the type of content and online sources crawled: The crawlers may be exposed to a wide variety of content and online sources, including most forms of publicly available online data. Type of modality covered: Text, Image Summary of the most relevant domain names crawled: The portion of the model's training corpus derived from data crawled and scraped from online sources includes technical documentation, open-source software, predominantly text-based reference sites, document sharing sites, and math sites. Top-level domains such as .com, .org, and .net are included alongside sites from a range of different countries. 2.4 User data Was data from user interactions with the AI model (e.g. user input and prompts) used to train the model? Yes Was data collected from user interactions with the provider’s other services or products used to train the model? Yes If yes, provide a general description of the provider’s services or products that were used to collect the user data: To the extent permitted by Anthropic’s terms of service, privacy policy, and other contracts, and in line with applicable law, if a user explicitly reports feedback or bugs to us (e.g., via thumbs and feedback buttons) or otherwise chooses to allow us to use their data, then chats and coding session data may be used in model training. More information is available at https://privacy.claude.com/en/articles/7996868-is-my-data-used-for-model-training We may also incorporate data derived from Anthropic employees’ use of internal-only model versions. 5 Training Data Summary Type of modality covered: Text 2.5 Synthetic data Was synthetic AI-generated data created by the provider or on their behalf to train the model? Yes If yes, modality of the synthetic data: Text, Image If yes, specify the general-purpose AI model(s) used to generate the synthetic data if available on the market: Synthetic data was provided by speech to text models, large language models (LLM), and vision-language models (VLM). Information about other AI models, including provider’s own AI model(s) not available on the market, used to generate synthetic data to train the model to which this Summary applies: Some synthetic data used in training was generated by Anthropic models not available on the market. 2.6 Other sources of data Have data sources other than those described in Sections 2.1 to 2.5 been used to train the model? Yes If yes, provide a narrative description of these data sources and the data: A portion of the data corpus comes from acquired physical texts. 1. Data processing aspects 3. Data processing aspects 3.1. Respect of reservation of rights from text and data mining exception or limitation 6 Training Data Summary Are you a Signatory to the Code of Practice for general-purpose AI models that includes commitments to respect reservations of rights from the TDM exception or limitation? Yes Describe the measures implemented before model training to respect reservations of rights from the TDM exception or limitation before and during data collection, including the opt-out protocols and solutions honoured by the provider or, as applicable, by third parties from which datasets have been obtained: ClaudeBot respects crawl-delay and disallow directives in robots.txt files where appropriate, as well as anti-circumvention technologies such as paywalls, password protection, and CAPTCHAs. More information on our crawlers and how they work can be found at: https://support.claude.com/en/articles/8896518-does-anthropic-crawl-data-from-the-web-and-how-can-site-owners-block-the-crawler. 3.2 Removal of illegal content General description of measures taken: We take a number of protective measures to remove illegal content from the training corpus such as active filtering, scoring, moderation, and blocking. 3.3. Other information (optional) Other relevant information about data processing (optional): Not applicable 7 ===== _Claude Mythos Preview Training Data Summary .pdf ===== Public Summary of Training Content Claude Mythos Preview Training Data Summary Version of the Summary: Version #1 Last update: July 24, 2026 General information 1. General information 1.1. Provider identification Provider name and contact details: Anthropic Ireland, Limited 6th Floor South Bank House, Barrow Street, Dublin 4, Dublin Ireland Authorised representative name and contact details: Not applicable 1.2. Model identification Versioned model name(s): Claude Mythos Preview Model Card: www.anthropic.com/system-cards Model dependencies: Not applicable Date of placement of the model on the Union market: June 2, 2026 1.3 Modalities, overall training data size and other characteristics Modality Select the modalities present in the training data, to the extent that they are identifiable Training data size For each selected modality, select the range within which the estimated total training data size for that modality falls. Dynamic datasets may be excluded from the estimation. Types of content For each selected modality, provide a general description of the type of content that has been included in the training data. 2 Training Data Summary Text ☐ Less than 1 billion tokens ☐ 1billion to 10 trillions tokens X More than 10 trillions tokens The training corpus for the model includes an array of text types, including short and long-form texts, software code, synthetic text, prose in a variety of languages, mathematical data, and prompts and preference data used during reinforcement learning. Image ☐ Less than 1 million images ☐ 1Million to1 billion images X More than 1 billion images The training corpus for the model includes an array of image types, including photographs, interleaved text and images from websites, computer graphics, and prompts and preference data used during reinforcement learning. Video Not applicable Audio Not applicable Other Not applicable Latest date of data acquisition/collection for model training: A number of different datasets, with varying publication and cut-off dates, are included in the training corpus, with some data being acquired/collected up to February 2026. Description of the linguistic characteristics of the overall training data: Training sources deliberately include a diverse range of global languages, both European and non-European, including those with relatively high numbers of speakers ( e.g., English, Chinese, French, Spanish) as well as comparably lower volumes ( e.g., Basque, Breton, Korean). 2. List of data sources 2. List of data sources 2.1. Publicly available datasets Have you used publicly available datasets to train the model? Yes If yes, specify the modality(ies) of the content covered by the datasets concerned: Text, Image List of large publicly available datasets: The training corpus is derived from several publicly accessible repositories, notably including Common Crawl, a repository of web crawl data, as well as specialized datasets available through platforms like GitHub and HuggingFace. 3 Training Data Summary General description of other publicly available datasets not listed above: Data from other publicly available datasets is included in the training corpus, including mathematical data and some image and caption datasets. 2.2 Private non-publicly available datasets obtained from third parties 2.2.1. Datasets commercially licensed by rightsholders or their representatives Have you concluded transactional commercial licensing agreement(s) with rightsholder(s) or with their representatives? Yes If yes, specify the modality(ies) of the content covered by the datasets concerned: Text, Image 2.2.2. Private datasets obtained from other third parties Have you obtained private datasets from third parties that are not licensed as described in Section 2.2.1, such as data obtained from providers of private databases, or data intermediaries? Yes If yes, specify the modality(ies) of the content covered by the datasets concerned: Text, Image If publicly known, list private datasets obtained from other third parties: Not applicable General description of non-publicly known private datasets obtained from third parties We obtain non-publicly known private datasets from third parties covering diverse domains and content types. 2.3 Data crawled and scraped from online sources Were crawlers used by the provider or on behalf of? Yes If yes, specify crawler name(s)/identifier(s): ClaudeBot Purposes of the crawler(s): ClaudeBot collects web content that could potentially contribute to the model’s training. 4 Training Data Summary General description of crawler behaviour: We aim to minimize disruption to website owners and be thoughtful about how quickly ClaudeBot crawls domains, including by respecting crawl-delay and disallow directives in robots.txt files where appropriate, and respecting anti-circumvention technologies such as paywalls, password protection, and CAPTCHAs. More information on our crawlers and how they work can be found at: https://support.claude.com/en/articles/8896518-does-anthropic-crawl-data-from-the-web-and-how-can-site-owners-block-the-crawler. Period of data collection: March 2024 - March 2026 Comprehensive description of the type of content and online sources crawled: The crawlers may be exposed to a wide variety of content and online sources, including most forms of publicly available online data. Type of modality covered: Text, Image Summary of the most relevant domain names crawled: The portion of the model's training corpus derived from data crawled and scraped from online sources includes technical documentation, open-source software, predominantly text-based reference sites, document sharing sites, and math sites. Top-level domains such as .com, .org, and .net are included alongside sites from a range of different countries. 2.4 User data Was data from user interactions with the AI model (e.g. user input and prompts) used to train the model? Yes Was data collected from user interactions with the provider’s other services or products used to train the model? Yes If yes, provide a general description of the provider’s services or products that were used to collect the user data: To the extent permitted by Anthropic’s terms of service, privacy policy, and other contracts, and in line with applicable law, if a user explicitly reports feedback or bugs to us (e.g., via thumbs and feedback buttons) or otherwise chooses to allow us to use their data, then chats and coding session data may be used in model training. More information is available at https://privacy.claude.com/en/articles/7996868-is-my-data-used-for-model-training We may also incorporate data derived from 5 Training Data Summary Anthropic employees’ use of internal-only model versions. Type of modality covered: Text 2.5 Synthetic data Was synthetic AI-generated data created by the provider or on their behalf to train the model? Yes If yes, modality of the synthetic data: Text, Image If yes, specify the general-purpose AI model(s) used to generate the synthetic data if available on the market: Synthetic data was provided by speech to text models, large language models (LLM), and vision-language models (VLM). Information about other AI models, including provider’s own AI model(s) not available on the market, used to generate synthetic data to train the model to which this Summary applies: Some synthetic data used in training was generated by Anthropic models not available on the market. 2.6 Other sources of data Have data sources other than those described in Sections 2.1 to 2.5 been used to train the model? Yes If yes, provide a narrative description of these data sources and the data: A portion of the data corpus comes from acquired physical texts. 6 Training Data Summary 1. 3. Data processing aspects 3.1. Respect of reservation of rights from text and data mining exception or limitation Are you a Signatory to the Code of Practice for general-purpose AI models that includes commitments to respect reservations of rights from the TDM exception or limitation? Yes Describe the measures implemented before model training to respect reservations of rights from the TDM exception or limitation before and during data collection, including the opt-out protocols and solutions honoured by the provider or, as applicable, by third parties from which datasets have been obtained: ClaudeBot respects crawl-delay and disallow directives in robots.txt files where appropriate, as well as anti-circumvention technologies such as paywalls, password protection, and CAPTCHAs. More information on our crawlers and how they work can be found at: https://support.claude.com/en/articles/8896518-does-anthropic-crawl-data-from-the-web-and-how-can-site-owners-block-the-crawler. 3.2 Removal of illegal content General description of measures taken: We take a number of protective measures to remove illegal content from the training corpus such as active filtering, scoring, moderation, and blocking. 3.3. Other information (optional) Other relevant information about data processing (optional): Not applicable. 7