GPAI Ledger The public record of EU AI Act training-data summaries

GPAI LedgerAnthropic trust-center bundle (Anthropic) › Capture 19 Aug 2026

Anthropic trust-center bundle — capture 20260819T202531Z

ProviderAnthropic
Targetprovider site — https://trust.anthropic.com/doc/trust-zip?r=… (signed URL; token masked, not linked)
Fetched (UTC)2026-08-17T10:01:23Z
Stored file4d2f4ba23f6af556bd0fad6cbc2fa40c1fad17fc98d40f4834d06ec32bd3aa14.zip (1,322,420 bytes)
SHA-2564d2f4ba23f6af556bd0fad6cbc2fa40c1fad17fc98d40f4834d06ec32bd3aa14
OpenTimestamps proof4d2f4ba23f6af556bd0fad6cbc2fa40c1fad17fc98d40f4834d06ec32bd3aa14.zip.ots (calendar-attested; anchored in bitcoin over time)
Waybacknot saved
Prior capture of this target20c0f32b59838a3966576ac2e074b5742f5cc256ee6374ce3f0a5441774cf647 (that capture was pruned as content-identical noise — by the prune rule its content is identical to a retained version of this target; its hash is in the event log)
NotesSCOPE REPACK (19 Aug 2026): this capture contains only the Art. 53 training-data summary documents from Anthropic's public trust-center bulk download. The full bundle also contained documents marked confidential/no-redistribution (SOC report, architecture overviews, certificates); those are NOT stored or served — their existence is recorded below by filename and SHA-256 only.

Verify: sha256sum 4d2f4ba23f6af556bd0fad6cbc2fa40c1fad17fc98d40f4834d06ec32bd3aa14.zip must equal the hash above (the filename IS the expected hash); ots verify 4d2f4ba23f6af556bd0fad6cbc2fa40c1fad17fc98d40f4834d06ec32bd3aa14.zip.ots -f 4d2f4ba23f6af556bd0fad6cbc2fa40c1fad17fc98d40f4834d06ec32bd3aa14.zip (opentimestamps.org) proves the capture time (fresh proofs report 'pending' until bitcoin-anchored, typically within a day).

Files in this bundle

FileSHA-256
Claude Mythos 5 and Claude Fable 5 Training Data Summary .pdf765afb9d4e62fa40…Art. 53 summary
Claude Opus 4.7 Training Data Summary .pdf184c573a50ec4e70…Art. 53 summary
Claude Opus 4.8 Training Data Summary .pdfccac79227fc4e9fd…Art. 53 summary
Claude Opus 5 Training Data Summary .pdf06dfc855a9905ffd…Art. 53 summary
Claude Sonnet 5 Training Data Summary .pdfc84c83a1215c56cf…Art. 53 summary
_Claude Mythos Preview Training Data Summary .pdf038c305c4760c4d2…Art. 53 summary

Extracted text (Art. 53 summaries)

Machine-extracted text (layout may be lost; the authoritative content is the stored file above).

Files recorded but not stored

These bundle members are outside Art. 53 scope (some carry confidentiality markings); they are recorded by name and SHA-256 so their identity stays provable, but their bytes are not archived or served.

FileSHA-256
AB 2013 Training Data Documentation [June 2026].pdfdf55f60c9bc1cd11…
ACR for Claude Android Enterprise - May 2026 - Anthropic PBC.pdf541652aea109b82d…
ACR for Claude Web Enterprise - April 2026 - Anthropic PBC.pdfb428badebb1568f7…
ACR for Claude iOS Enterprise - May 2026 - Anthropic PBC.pdf35f9ea604acc09f8…
Anthropic Statement on Modern Slavery Act 2015.pdfd1d71f9119098a4e…
CVE-2026-22561 - DLL Search Order Hijacking in Claude for Windows installer.pdf50b69a0273d34256…
Claude Desktop 3P Security Overview.pdfa3fcdcc61d5802de…
Claude Mythos 5 + Fable 5 Model Documentation Form v2.pdfd4c04c6e2beb0f58…
Claude Mythos Preview Model Documentation Form v2 [PDF].pdf9fb3d243e0955890…
Claude Opus 4.7 Model Documentation Form for Downstream Providers_v2.pdf2bebf9186e1b22fd…
Claude Opus 4.8 Model Documentation Form for downstream providers_v2 (1).pdf6645a5afcaad0eb0…
Claude Opus 5_ Model Documentation Form for downstream providers [PDF].pdf5a740095beed9ef0…
Claude Sonnet 5 Model Documentation Form v2.pdfdf248a415f6f2b03…
Claude in Excel & PowerPoint Architecture Overview.pdf8dc1b3025d9368e5…
Frontier Compliance Framework_July 2026.pdf8e4d91e12861218e…
Office Agents 3P Architecture.pdf21214bde3519c4d3…
[Anthropic Ireland Limited] Cyber Essentials Certificate (2025).pdf223c71464e409116…
[Anthropic] 2025 Type 2 SOC 3 Report.pdf4a2eec78000029cc…
[Anthropic] Anthropic's Enterprise Security Posture.pdfae3a40f6cd75a651…
[Anthropic] Data Handling & Key Management.pdf4d86c308354881f8…
[Anthropic] Data Loss Prevention & Content Controls.pdfcd85d044e4ccf41f…
[Anthropic] ISO 27001 Certificate (2025).pdfb8d621cc3c9ac880…
[Anthropic] ISO 42001 Certificate (2025).pdfa72bbe24b44b5c17…
[Anthropic] Identity & Access Controls.pdf1b9f0c3575e892b5…
[Anthropic] Isolation & Connectivity.pdf092b7815b2e4ca76…
[Anthropic] Security and Privacy Design of Anthropic Data Retention and Review.pdfa6aaa6952308e11c…
v1.0 Claude Code FISMA Best Practices.pdf1c4ab4052647fe75…
===== Claude Mythos 5 and Claude Fable 5 Training Data Summary .pdf =====
Public  Summary  of  Training
Content

Claude

Mythos

5

&

Claude

Fable

5

Training  Data  Summary   Version  of  the   Summary:   Version  #1   Last  update:
July

24,

2026

1.  General  information
1.1.

Provider

identification

Provider  name  and  contact  details:
Anthropic  Ireland,  Limited  6th  Floor  South  Bank  House,  Barrow  Street,  Dublin  4,  Dublin  Ireland

Authorised  representative  name  and  contact  details:  Not  applicable

1.2.  Model  identification
Versioned  model  name(s):
 Claude  Mythos  5  and  Claude  Fable  5  Model  Card:  www.anthropic.com/system-cards
Model  dependencies:  Not  applicable
  Date  of  placement  of  the  model  on  the  Union  market:  June  9,  2026
1.3  Modalities,  overall  training  data  size  and  other  characteristics  Modality  Select  the  modalities  present  in  the  training  data,  to  the  extent  that  they  are
identifiable

Training  data  size  For  each  selected  modality,  select  the  range  within  which  the  estimated  total  training  data  size  for  that  modality  falls.  Dynamic  datasets  may  be  excluded  from  the  estimation.
Types  of  content  For  each  selected  modality,  provide  a  general  description  of  the  type  of  content  that  has  been  included  in  the
training

data.

Text

☐

Less

than

1

billion

tokens
 ☐  1billion  to  10  trillions  tokens  X  More  than  10  trillions  tokens
The  training  corpus  for  the  model  includes  an  array  of  text  types,  including  short  and  long-form  texts,  software  code,  synthetic  text,  prose  in  a  variety  of  languages,  mathematical  data,  and  prompts  and  preference  data  used  during  reinforcement  learning.
Image
☐

Less

than

1

million

images

☐

1Million

to1

billion

images
 X  More  than  1  billion  images
The  training  corpus  for  the  model  includes  an  array  of  image  types,  including  photographs,  interleaved
2

Training  Data  Summary    text  and  images  from  websites,  computer  graphics,  and  prompts  and  preference  data  used  during  reinforcement  learning.   Video  Not  applicable  Audio  Not  applicable  Other  Not  applicable
 Latest  date  of  data  acquisition/collection  for  model  training:
A  number  of  different  datasets,  with  varying  publication  and  cut-off  dates,  are  included  in  the  training  corpus,  with  some  data  being  acquired/collected  up  to  April  2026.
Description  of  the  linguistic  characteristics  of  the  overall  training  data:
Training  sources  deliberately  include  a  diverse  range  of  global  languages,  both  European  and  non-European,  including  those  with  relatively  high  numbers  of  speakers  ( e.g.,  English,  Chinese,  French,  Spanish)  as  well  as  comparably  lower  volumes  ( e.g.,  Basque,  Breton,  Korean).
2.  List  of  data  sources
2.  List  of  data  sources
2.1.  Publicly  available  datasets
Have  you  used  publicly  available  datasets  to  train  the  model?    Yes
If  yes,  specify  the  modality(ies)  of  the  content  covered  by  the  datasets  concerned:
Text,  Image
List  of  large  publicly  available  datasets:
The  training  corpus  is  derived  from  several  publicly  accessible  repositories,  notably  including  Common  Crawl,  a  repository  of  web  crawl  data,  as  well  as  specialized  datasets  available  through  platforms  like  GitHub  and  HuggingFace.
General  description  of  other  publicly  available  datasets  not  listed  above:
Data  from  other  publicly  available  datasets  is  included  in  the  training  corpus,  including  mathematical  data  and  some  image  and  caption  datasets.

2.2  Private  non-publicly  available  datasets  obtained  from  third
parties
  2.2.1.  Datasets  commercially  licensed  by  rightsholders  or  their  representatives  Have  you  concluded  transactional  commercial  licensing  agreement(s)  with  rightsholder(s)  or  with  their  representatives?
Yes
3

Training  Data  Summary   If  yes,  specify  the  modality(ies)  of  the  content  covered  by  the  datasets  concerned:   Text,  Image
2.2.2.  Private  datasets  obtained  from  other  third  parties  Have  you  obtained  private  datasets  from  third  parties  that  are  not  licensed  as  described  in  Section  2.2.1,  such  as  data  obtained  from  providers  of  private  databases,  or  data  intermediaries?
Yes
If  yes,  specify  the  modality(ies)  of  the  content  covered  by  the  datasets  concerned:
Text,  Image

If  publicly  known,  list  private  datasets  obtained  from  other  third  parties:   Not  applicable

General  description  of  non-publicly  known  private  datasets  obtained  from  third  parties
We  obtain  non-publicly  known  private  datasets  from  third  parties  covering  diverse  domains  and  content  types.

2.3  Data  crawled  and  scraped  from  online  sources
Were  crawlers  used  by  the  provider  or  on  behalf  of?

Yes

If  yes,  specify  crawler  name(s)/identifier(s):
ClaudeBot

Purposes  of  the  crawler(s):
ClaudeBot  collects  web  content  that  could  potentially  contribute  to  the  model’s  training.
General  description  of  crawler  behaviour:
We  aim  to  minimize  disruption  to  website  owners  and  be  thoughtful  about  how  quickly  ClaudeBot  crawls  domains,    including  by  respecting  crawl-delay  and  disallow  directives  in  robots.txt  files,  where  appropriate,  and  respecting  anti-circumvention  technologies  such  as  paywalls,  password  protection,  and  CAPTCHAs.  More  information  on  our  crawlers  and  how  they  work  can  be  found  at:  https://support.claude.com/en/articles/8896518-does-anthropic-crawl-data-from-the-web-and-how-can-site-owners-block-the-crawler.
Period  of  data  collection:  March  2024  -  April  2026
Comprehensive  description  of  the  type  of  content  and  online  sources  crawled:
The  crawlers  may  be  exposed  to  a  wide  variety  of  content  and  online  sources,  including  most  forms  of  publicly  available  online  data.
4

Training  Data  Summary   Type  of  modality  covered:  Text,  Image

Summary  of  the  most  relevant  domain  names  crawled:
The  portion  of  the  model's  training  corpus  derived  from  data  crawled  and  scraped  from  online  sources  includes  technical  documentation,  open-source  software,  predominantly  text-based  reference  sites,  document  sharing  sites,  and  math  sites.  Top-level  domains  such  as  .com,  .org,  and  .net  are  included  alongside  sites  from  a  range  of  different  countries.
2.4  User  data   Was  data  from  user  interactions  with  the  AI  model  (e.g.  user  input  and  prompts)  used  to  train  the  model?
 Yes

Was  data  collected  from  user  interactions  with  the  provider’s  other  services  or  products  used  to  train  the  model?
 Yes
If  yes,  provide  a  general  description  of  the  provider’s  services  or  products  that  were  used  to  collect  the  user  data:
 To  the  extent  permitted  by  Anthropic’s  terms  of  service,  privacy  policy,  and  other  contracts,  and  in  line  with  applicable  law,  if  a  user  explicitly  reports  feedback  or  bugs  to  us  (e.g.,  via  thumbs  and  feedback  buttons)  or  otherwise  chooses  to  allow  us  to  use  their  data,  then  chats  and  coding  session  data  may  be  used  in  model  training.  More  information  is  available  at  https://privacy.claude.com/en/articles/7996868-is-my-data-used-for-model-training    We  may  also  incorporate  data  derived  from  Anthropic  employees’  use  of  internal-only  model  versions.
Type  of  modality  covered:  Text

2.5  Synthetic  data   Was  synthetic  AI-generated  data  created  by  the  provider  or  on  their  behalf  to  train  the  model?

 Yes
If  yes,  modality  of  the  synthetic  data:

Text,  Image
5

Training  Data  Summary   If  yes,  specify  the  general-purpose  AI  model(s)  used  to  generate  the  synthetic  data  if  available  on  the  market:

 Synthetic  data  was  provided  by  speech  to  text  models,  large  language  models  (LLM),  and  vision-language  models  (VLM).
 Information  about  other  AI  models,  including  provider’s  own  AI  model(s)  not  available  on  the  market,  used  to  generate  synthetic  data  to  train  the  model  to  which  this  Summary  applies:

 Some  synthetic  data  used  in  training  was  generated  by  Anthropic  models  not  available  on  the  market.

2.6  Other  sources  of  data   Have  data  sources  other  than  those  described  in  Sections  2.1  to  2.5  been  used  to  train  the  model?
Yes

If  yes,  provide  a  narrative  description  of  these  data  sources  and  the  data:
 A  portion  of  the  data  corpus  comes  from  acquired  physical  texts.

1.  Data  processing  aspects
3.  Data  processing  aspects
3.1.

Respect

of

reservation

of

rights

from

text

and

data

mining

exception

or

limitation

  Are  you  a  Signatory  to  the  Code  of  Practice  for  general-purpose  AI  models  that  includes  commitments  to  respect  reservations  of  rights  from  the  TDM  exception  or  limitation?

Yes

Describe  the  measures  implemented  before  model  training  to  respect  reservations  of  rights  from  the  TDM  exception  or  limitation  before  and  during  data  collection,  including  the  opt-out  protocols  and  solutions  honoured  by  the  provider  or,  as  applicable,  by  third  parties  from  which  datasets  have  been  obtained:
 ClaudeBot  respects  crawl-delay  and  disallow  directives  in  robots.txt  files  where  appropriate,  as  well  as  anti-circumvention  technologies  such  as  paywalls,  password  protection,  and  CAPTCHAs.  More  information  on  our  crawlers  and  how  they  work  can  be
6

Training  Data  Summary   found  at:  https://support.claude.com/en/articles/8896518-does-anthropic-crawl-data-from-the-web-and-how-can-site-owners-block-the-crawler.

3.2  Removal  of  illegal  content
General  description  of  measures  taken:
We  take  a  number  of  protective  measures  to  remove  illegal  content  from  the  training  corpus  such  as  active  filtering,  scoring,  moderation,  and  blocking.

3.3.  Other  information  (optional)  Other  relevant  information  about  data  processing  (optional):
Not  applicable

7

===== Claude Opus 4.7 Training Data Summary .pdf =====
Public  Summary  of  Training
Content

Claude

Opus

4.7

Training  Data  Summary   Version  of  the  Summary:   Version  #1   Last  update:
July

24,

2026
 General  information

1.  General  information
1.1.

Provider

identification

Provider  name  and  contact  details:
Anthropic  Ireland,  Limited  6th  Floor  South  Bank  House,  Barrow  Street,  Dublin  4,  Dublin  Ireland

Authorised  representative  name  and  contact  details:  Not  applicable

1.2.  Model  identification
Versioned  model  name(s):
 Claude  Opus  4.7  Model  Card:  www.anthropic.com/system-cards
Model  dependencies:  Not  applicable
  Date  of  placement  of  the  model  on  the  Union  market:  April  16,  2026
1.3  Modalities,  overall  training  data  size  and  other  characteristics  Modality  Select  the  modalities  present  in  the  training  data,  to  the  extent  that  they  are
identifiable

Training  data  size  For  each  selected  modality,  select  the  range  within  which  the  estimated  total  training  data  size  for  that  modality  falls.  Dynamic  datasets  may  be  excluded  from  the  estimation.
Types  of  content  For  each  selected  modality,  provide  a  general  description  of  the  type  of  content  that  has  been  included  in  the
training

data.

Text

☐

Less

than

1

billion

tokens
 ☐  1billion  to  10  trillions  tokens  X  More  than  10  trillions  tokens
The  training  corpus  for  the  model  includes  an  array  of  text  types,  including  short  and  long-form  texts,  software  code,  synthetic  text,  prose  in  a  variety  of  languages,  mathematical  data,  and  prompts  and  preference  data  used  during  reinforcement  learning.
2

Training  Data  Summary
Image
☐

Less

than

1

million

images

☐

1Million

to1

billion

images
 X  More  than  1  billion  images
The  training  corpus  for  the  model  includes  an  array  of  image  types,  including  photographs,  interleaved  text  and  images  from  websites,  computer  graphics,  and  prompts  and  preference  data  used  during  reinforcement  learning.   Video  Not  applicable  Audio  Not  applicable  Other  Not  applicable
 Latest  date  of  data  acquisition/collection  for  model  training:
A  number  of  different  datasets,  with  varying  publication  and  cut-off  dates,  are  included  in  the  training  corpus,  with  some  data  being  acquired/collected  up  to  April  2026.
Description  of  the  linguistic  characteristics  of  the  overall  training  data:
Training  sources  deliberately  include  a  diverse  range  of  global  languages,  both  European  and  non-European,  including  those  with  relatively  high  numbers  of  speakers  ( e.g.,  English,  Chinese,  French,  Spanish)  as  well  as  comparably  lower  volumes  ( e.g.,  Basque,  Breton,  Korean).
2.  List  of  data  sources
2.  List  of  data  sources
2.1.  Publicly  available  datasets
Have  you  used  publicly  available  datasets  to  train  the  model?    Yes
If  yes,  specify  the  modality(ies)  of  the  content  covered  by  the  datasets  concerned:
Text,  Image
List  of  large  publicly  available  datasets:
The  training  corpus  is  derived  from  several  publicly  accessible  repositories,  notably  including  Common  Crawl,  a  repository  of  web  crawl  data,  as  well  as  specialized  datasets  available  through  platforms  like  GitHub  and  HuggingFace.
General  description  of  other  publicly  available  datasets  not  listed  above:
Data  from  other  publicly  available  datasets  is  included  in  the  training  corpus,  including  mathematical  data  and  some  image  and  caption  datasets.

2.2  Private  non-publicly  available  datasets  obtained  from  third
parties

3

Training  Data  Summary
2.2.1.  Datasets  commercially  licensed  by  rightsholders  or  their  representatives  Have  you  concluded  transactional  commercial  licensing  agreement(s)  with  rightsholder(s)  or  with  their  representatives?
Yes
If  yes,  specify  the  modality(ies)  of  the  content  covered  by  the  datasets  concerned:   Text,  Image
2.2.2.  Private  datasets  obtained  from  other  third  parties  Have  you  obtained  private  datasets  from  third  parties  that  are  not  licensed  as  described  in  Section  2.2.1,  such  as  data  obtained  from  providers  of  private  databases,  or  data  intermediaries?
Yes
If  yes,  specify  the  modality(ies)  of  the  content  covered  by  the  datasets  concerned:
Text,  Image

If  publicly  known,  list  private  datasets  obtained  from  other  third  parties:   Not  applicable

General  description  of  non-publicly  known  private  datasets  obtained  from  third  parties
We  obtain  non-publicly  known  private  datasets  from  third  parties  covering  diverse  domains  and  content  types.

2.3  Data  crawled  and  scraped  from  online  sources
Were  crawlers  used  by  the  provider  or  on  behalf  of?

Yes

If  yes,  specify  crawler  name(s)/identifier(s):
ClaudeBot

Purposes  of  the  crawler(s):
ClaudeBot  collects  web  content  that  could  potentially  contribute  to  the  model’s  training.
General  description  of  crawler  behaviour:
We  aim  to  minimize  disruption  to  website  owners  and  be  thoughtful  about  how  quickly  ClaudeBot  crawls  domains,    including  by  respecting  crawl-delay  and  disallow  directives  in  robots.txt  files,  where  appropriate,  and  respecting  anti-circumvention  technologies  such  as  paywalls,  password  protection,  and  CAPTCHAs.  More  information  on  our  crawlers  and  how  they  work  can  be  found  at:  https://support.claude.com/en/articles/8896518-does-anthropic-crawl-data-from-the-web-and-how-can-site-owners-block-the-crawler.
4

Training  Data  Summary   Period  of  data  collection:  March  2024  -  April  2026
Comprehensive  description  of  the  type  of  content  and  online  sources  crawled:
The  crawlers  may  be  exposed  to  a  wide  variety  of  content  and  online  sources,  including  most  forms  of  publicly  available  online  data.
Type  of  modality  covered:  Text,  Image

Summary  of  the  most  relevant  domain  names  crawled:
The  portion  of  the  model's  training  corpus  derived  from  data  crawled  and  scraped  from  online  sources  includes  technical  documentation,  open-source  software,  predominantly  text-based  reference  sites,  document  sharing  sites,  and  math  sites.  Top-level  domains  such  as  .com,  .org,  and  .net  are  included  alongside  sites  from  a  range  of  different  countries.
2.4  User  data   Was  data  from  user  interactions  with  the  AI  model  (e.g.  user  input  and  prompts)  used  to  train  the  model?
 Yes

Was  data  collected  from  user  interactions  with  the  provider’s  other  services  or  products  used  to  train  the  model?
 Yes
If  yes,  provide  a  general  description  of  the  provider’s  services  or  products  that  were  used  to  collect  the  user  data:
 To  the  extent  permitted  by  Anthropic’s  terms  of  service,  privacy  policy,  and  other  contracts,  and  in  line  with  applicable  law,  if  a  user  explicitly  reports  feedback  or  bugs  to  us  (e.g.,  via  thumbs  and  feedback  buttons)  or  otherwise  chooses  to  allow  us  to  use  their  data,  then  chats  and  coding  session  data  may  be  used  in  model  training.  More  information  is  available  at  https://privacy.claude.com/en/articles/7996868-is-my-data-used-for-model-training    We  may  also  incorporate  data  derived  from  Anthropic  employees’  use  of  internal-only  model  versions.
Type  of  modality  covered:  Text

2.5  Synthetic  data   Was  synthetic  AI-generated  data  created  by  the

 Yes
5

Training  Data  Summary   provider  or  on  their  behalf  to  train  the  model?
If  yes,  modality  of  the  synthetic  data:

Text,  Image
If  yes,  specify  the  general-purpose  AI  model(s)  used  to  generate  the  synthetic  data  if  available  on  the  market:

 Synthetic  data  was  provided  by  speech  to  text  models,  large  language  models  (LLM),  and  vision-language  models  (VLM).
 Information  about  other  AI  models,  including  provider’s  own  AI  model(s)  not  available  on  the  market,  used  to  generate  synthetic  data  to  train  the  model  to  which  this  Summary  applies:

 Some  synthetic  data  used  in  training  was  generated  by  Anthropic  models  not  available  on  the  market.

2.6  Other  sources  of  data   Have  data  sources  other  than  those  described  in  Sections  2.1  to  2.5  been  used  to  train  the  model?
Yes

If  yes,  provide  a  narrative  description  of  these  data  sources  and  the  data:
 A  portion  of  the  data  corpus  comes  from  acquired  physical  texts.

1.  Data  processing  aspects
3.  Data  processing  aspects
3.1.

Respect

of

reservation

of

rights

from

text

and

data

mining

exception

or

limitation

  Are  you  a  Signatory  to  the  Code  of  Practice  for  general-purpose  AI  models  that  includes  commitments  to  respect  reservations  of  rights  from  the  TDM  exception  or  limitation?

Yes

6

Training  Data  Summary
Describe  the  measures  implemented  before  model  training  to  respect  reservations  of  rights  from  the  TDM  exception  or  limitation  before  and  during  data  collection,  including  the  opt-out  protocols  and  solutions  honoured  by  the  provider  or,  as  applicable,  by  third  parties  from  which  datasets  have  been  obtained:
 ClaudeBot  respects  crawl-delay  and  disallow  directives  in  robots.txt  files  where  appropriate,  as  well  as  anti-circumvention  technologies  such  as  paywalls,  password  protection,  and  CAPTCHAs.  More  information  on  our  crawlers  and  how  they  work  can  be  found  at:  https://support.claude.com/en/articles/8896518-does-anthropic-crawl-data-from-the-web-and-how-can-site-owners-block-the-crawler.

3.2  Removal  of  illegal  content
General  description  of  measures  taken:
We  take  a  number  of  protective  measures  to  remove  illegal  content  from  the  training  corpus  such  as  active  filtering,  scoring,  moderation,  and  blocking.

3.3.  Other  information  (optional)  Other  relevant  information  about  data  processing  (optional):
Not  applicable

7

===== Claude Opus 4.8 Training Data Summary .pdf =====
Public  Summary  of  Training
Content

Claude

Opus

4.8

Training  Data  Summary   Version  of  the  Summary:   Version  #1   Last  update:
July

24,

2026
 General  information

1.  General  information
1.1.

Provider

identification

Provider  name  and  contact  details:
Anthropic  Ireland,  Limited  6th  Floor  South  Bank  House,  Barrow  Street,  Dublin  4,  Dublin  Ireland

Authorised  representative  name  and  contact  details:  Not  applicable

1.2.  Model  identification
Versioned  model  name(s):
 Claude  Opus  4.8  Model  Card:  www.anthropic.com/system-cards
Model  dependencies:  Not  applicable
  Date  of  placement  of  the  model  on  the  Union  market:  May  28,  2026
1.3  Modalities,  overall  training  data  size  and  other  characteristics  Modality  Select  the  modalities  present  in  the  training  data,  to  the  extent  that  they  are
identifiable

Training  data  size  For  each  selected  modality,  select  the  range  within  which  the  estimated  total  training  data  size  for  that  modality  falls.  Dynamic  datasets  may  be  excluded  from  the  estimation.
Types  of  content  For  each  selected  modality,  provide  a  general  description  of  the  type  of  content  that  has  been  included  in  the
training

data.

Text

☐

Less

than

1

billion

tokens
 ☐  1billion  to  10  trillions  tokens  X  More  than  10  trillions  tokens
The  training  corpus  for  the  model  includes  an  array  of  text  types,  including  short  and  long-form  texts,  software  code,  synthetic  text,  prose  in  a  variety  of  languages,  mathematical  data,  and  prompts  and  preference  data  used  during  reinforcement  learning.    Image
☐

Less

than

1

million

images

☐

1Million

to1

billion

images

The  training  corpus  for  the  model  includes  an  array  of  image  types,
2

Training  Data  Summary   X  More  than  1  billion  images
including  photographs,  interleaved  text  and  images  from  websites,  computer  graphics,  and  prompts  and  preference  data  used  during  reinforcement  learning.   Video  Not  applicable  Audio  Not  applicable  Other  Not  applicable
 Latest  date  of  data  acquisition/collection  for  model  training:
A  number  of  different  datasets,  with  varying  publication  and  cut-off  dates,  are  included  in  the  training  corpus,  with  some  data  being  acquired/collected  up  to  May  2026.
Description  of  the  linguistic  characteristics  of  the  overall  training  data:
Training  sources  deliberately  include  a  diverse  range  of  global  languages,  both  European  and  non-European,  including  those  with  relatively  high  numbers  of  speakers  ( e.g.,  English,  Chinese,  French,  Spanish)  as  well  as  comparably  lower  volumes  ( e.g.,  Basque,  Breton,  Korean).
2.  List  of  data  sources
2.  List  of  data  sources
2.1.  Publicly  available  datasets
Have  you  used  publicly  available  datasets  to  train  the  model?    Yes
If  yes,  specify  the  modality(ies)  of  the  content  covered  by  the  datasets  concerned:
Text,  Image
List  of  large  publicly  available  datasets:
The  training  corpus  is  derived  from  several  publicly  accessible  repositories,  notably  including  Common  Crawl,  a  repository  of  web  crawl  data,  as  well  as  specialized  datasets  available  through  platforms  like  GitHub  and  HuggingFace.
General  description  of  other  publicly  available  datasets  not  listed  above:
Data  from  other  publicly  available  datasets  is  included  in  the  training  corpus,  including  mathematical  data  and  some  image  and  caption  datasets.

2.2  Private  non-publicly  available  datasets  obtained  from  third
parties
  2.2.1.  Datasets  commercially  licensed  by  rightsholders  or  their  representatives  Have  you  concluded  transactional  commercial  licensing  agreement(s)  with  rightsholder(s)  or  with  their  representatives?
Yes
3

Training  Data  Summary   If  yes,  specify  the  modality(ies)  of  the  content  covered  by  the  datasets  concerned:   Text,  Image
2.2.2.  Private  datasets  obtained  from  other  third  parties  Have  you  obtained  private  datasets  from  third  parties  that  are  not  licensed  as  described  in  Section  2.2.1,  such  as  data  obtained  from  providers  of  private  databases,  or  data  intermediaries?
Yes
If  yes,  specify  the  modality(ies)  of  the  content  covered  by  the  datasets  concerned:
Text,  Image

If  publicly  known,  list  private  datasets  obtained  from  other  third  parties:   Not  applicable

General  description  of  non-publicly  known  private  datasets  obtained  from  third  parties
We  obtain  non-publicly  known  private  datasets  from  third  parties  covering  diverse  domains  and  content  types.

2.3  Data  crawled  and  scraped  from  online  sources
Were  crawlers  used  by  the  provider  or  on  behalf  of?

Yes

If  yes,  specify  crawler  name(s)/identifier(s):
ClaudeBot

Purposes  of  the  crawler(s):
ClaudeBot  collects  web  content  that  could  potentially  contribute  to  the  model’s  training.
General  description  of  crawler  behaviour:
We  aim  to  minimize  disruption  to  website  owners  and  be  thoughtful  about  how  quickly  ClaudeBot  crawls  domains,    including  by  respecting  crawl-delay  and  disallow  directives  in  robots.txt  files,  where  appropriate,  and  respecting  anti-circumvention  technologies  such  as  paywalls,  password  protection,  and  CAPTCHAs.  More  information  on  our  crawlers  and  how  they  work  can  be  found  at:  https://support.claude.com/en/articles/8896518-does-anthropic-crawl-data-from-the-web-and-how-can-site-owners-block-the-crawler.
Period  of  data  collection:  March  2024  -  May  2026
Comprehensive  description  of  the  type  of  content  and  online  sources  crawled:
The  crawlers  may  be  exposed  to  a  wide  variety  of  content  and  online  sources,  including  most  forms  of  publicly  available  online  data.
4

Training  Data  Summary   Type  of  modality  covered:  Text,  Image

Summary  of  the  most  relevant  domain  names  crawled:
The  portion  of  the  model's  training  corpus  derived  from  data  crawled  and  scraped  from  online  sources  includes  technical  documentation,  open-source  software,  predominantly  text-based  reference  sites,  document  sharing  sites,  and  math  sites.  Top-level  domains  such  as  .com,  .org,  and  .net  are  included  alongside  sites  from  a  range  of  different  countries.
2.4  User  data   Was  data  from  user  interactions  with  the  AI  model  (e.g.  user  input  and  prompts)  used  to  train  the  model?
 Yes

Was  data  collected  from  user  interactions  with  the  provider’s  other  services  or  products  used  to  train  the  model?
 Yes
If  yes,  provide  a  general  description  of  the  provider’s  services  or  products  that  were  used  to  collect  the  user  data:
 To  the  extent  permitted  by  Anthropic’s  terms  of  service,  privacy  policy,  and  other  contracts,  and  in  line  with  applicable  law,  if  a  user  explicitly  reports  feedback  or  bugs  to  us  (e.g.,  via  thumbs  and  feedback  buttons)  or  otherwise  chooses  to  allow  us  to  use  their  data,  then  chats  and  coding  session  data  may  be  used  in  model  training.  More  information  is  available  at  https://privacy.claude.com/en/articles/7996868-is-my-data-used-for-model-training    We  may  also  incorporate  data  derived  from  Anthropic  employees’  use  of  internal-only  model  versions.
Type  of  modality  covered:  Text

2.5  Synthetic  data   Was  synthetic  AI-generated  data  created  by  the  provider  or  on  their  behalf  to  train  the  model?

 Yes
If  yes,  modality  of  the  synthetic  data:

Text,  Image
5

Training  Data  Summary   If  yes,  specify  the  general-purpose  AI  model(s)  used  to  generate  the  synthetic  data  if  available  on  the  market:

 Synthetic  data  was  provided  by  speech  to  text  models,  large  language  models  (LLM),  and  vision-language  models  (VLM).
 Information  about  other  AI  models,  including  provider’s  own  AI  model(s)  not  available  on  the  market,  used  to  generate  synthetic  data  to  train  the  model  to  which  this  Summary  applies:

 Some  synthetic  data  used  in  training  was  generated  by  Anthropic  models  not  available  on  the  market.

2.6  Other  sources  of  data   Have  data  sources  other  than  those  described  in  Sections  2.1  to  2.5  been  used  to  train  the  model?
Yes

If  yes,  provide  a  narrative  description  of  these  data  sources  and  the  data:
 A  portion  of  the  data  corpus  comes  from  acquired  physical  texts.

1.  Data  processing  aspects
3.  Data  processing  aspects
3.1.

Respect

of

reservation

of

rights

from

text

and

data

mining

exception

or

limitation

  Are  you  a  Signatory  to  the  Code  of  Practice  for  general-purpose  AI  models  that  includes  commitments  to  respect  reservations  of  rights  from  the  TDM  exception  or  limitation?

Yes

Describe  the  measures  implemented  before  model  training  to  respect  reservations  of  rights  from  the  TDM  exception  or  limitation  before  and  during  data  collection,  including  the  opt-out  protocols  and  solutions  honoured  by  the  provider  or,  as  applicable,  by  third  parties  from  which  datasets  have  been  obtained:
 ClaudeBot  respects  crawl-delay  and  disallow  directives  in  robots.txt  files  where  appropriate,  as  well  as  anti-circumvention  technologies  such  as  paywalls,  password  protection,  and  CAPTCHAs.  More  information  on  our  crawlers  and  how  they  work  can  be
6

Training  Data  Summary   found  at:  https://support.claude.com/en/articles/8896518-does-anthropic-crawl-data-from-the-web-and-how-can-site-owners-block-the-crawler.

3.2  Removal  of  illegal  content
General  description  of  measures  taken:
We  take  a  number  of  protective  measures  to  remove  illegal  content  from  the  training  corpus  such  as  active  filtering,  scoring,  moderation,  and  blocking.

3.3.  Other  information  (optional)  Other  relevant  information  about  data  processing  (optional):
Not  applicable

7

===== Claude Opus 5 Training Data Summary .pdf =====
Public  Summary  of  Training
Content

Claude

Opus

5

Training  Data  Summary

 Version  of  the  Summary:   Version  #1
Last  update:
July

23,

2026
 General  information

1.  General  information   1.1.  Provider  identification
Provider  name  and  contact  details:
Anthropic  Ireland,  Limited  6th  Floor  South  Bank  House,  Barrow  Street,  Dublin  4,  Dublin  Ireland

Authorised  representative  name  and  contact  details:  Not  applicable

1.2.  Model  identification
Versioned  model  name(s):
 Claude  Opus  5  Model  Card:  www.anthropic.com/system-cards
Model  dependencies:  Not  applicable
  Date  of  placement  of  the  model  on  the  Union  market:  July  23,  2026
1.3  Modalities,  overall  training  data  size  and  other  characteristics  Modality  Select  the  modalities  present  in  the  training  data,  to  the  extent  that  they  are
identifiable

Training  data  size  For  each  selected  modality,  select  the  range  within  which  the  estimated  total  training  data  size  for  that  modality  falls.  Dynamic  datasets  may  be  excluded  from  the  estimation.
Types  of  content  For  each  selected  modality,  provide  a  general  description  of  the  type  of  content  that  has  been  included  in  the
training

data.

2

Training  Data  Summary
Text

☐

Less

than

1

billion

tokens
 ☐  1billion  to  10  trillions  tokens  X  More  than  10  trillions  tokens
The  training  corpus  for  the  model  includes  an  array  of  text  types,  including  short  and  long-form  texts,  software  code,  synthetic  text,  prose  in  a  variety  of  languages,  mathematical  data,  and  prompts  and  preference  data  used  during  reinforcement  learning.
Image
☐

Less

than

1

million

images

☐

1Million

to1

billion

images
 X  More  than  1  billion  images
The  training  corpus  for  the  model  includes  an  array  of  image  types,  including  photographs,  interleaved  text  and  images  from  websites,  computer  graphics,  and  prompts  and  preference  data  used  during  reinforcement  learning.   Video  Not  applicable   Audio  Not  applicable  Other  Not  applicable
 Latest  date  of  data  acquisition/collection  for  model  training:
A  number  of  different  datasets,  with  varying  publication  and  cut-off  dates,  are  included  in  the  training  corpus,  with  some  data  being  acquired/collected  up  to  July  2026.
Description  of  the  linguistic  characteristics  of  the  overall  training  data:
Training  sources  deliberately  include  a  diverse  range  of  global  languages,  both  European  and  non-European,  including  those  with  relatively  high  numbers  of  speakers  ( e.g.,  English,  Chinese,  French,  Spanish)  as  well  as  comparably  lower  volumes  ( e.g.,  Basque,  Breton,  Korean).
2.  List  of  data  sources
2.  List  of  data  sources
 2.1.  Publicly  available  datasets
Have  you  used  publicly  available  datasets  to  train  the  model?    Yes
If  yes,  specify  the  modality(ies)  of  the  content  covered  by  the  datasets  concerned:
Text,  Image
List  of  large  publicly  available  datasets:
The  training  corpus  is  derived  from  several  publicly  accessible  repositories,  notably  including  Common  Crawl,  a
3

Training  Data  Summary   repository  of  web  crawl  data,  as  well  as  specialized  datasets  available  through  platforms  like  GitHub  and  HuggingFace.
General  description  of  other  publicly  available  datasets  not  listed  above:
Data  from  other  publicly  available  datasets  is  included  in  the  training  corpus,  including  mathematical  data  and  some  image  and  caption  datasets.

  2.2  Private  non-publicly  available  datasets  obtained  from  third
parties
  2.2.1.  Datasets  commercially  licensed  by  rightsholders  or  their  representatives  Have  you  concluded  transactional  commercial  licensing  agreement(s)  with  rightsholder(s)  or  with  their  representatives?
Yes
If  yes,  specify  the  modality(ies)  of  the  content  covered  by  the  datasets  concerned:   Text,  Image
2.2.2.  Private  datasets  obtained  from  other  third  parties  Have  you  obtained  private  datasets  from  third  parties  that  are  not  licensed  as  described  in  Section  2.2.1,  such  as  data  obtained  from  providers  of  private  databases,  or  data  intermediaries?
Yes
If  yes,  specify  the  modality(ies)  of  the  content  covered  by  the  datasets  concerned:
Text,  Image

If  publicly  known,  list  private  datasets  obtained  from  other  third  parties:   Not  applicable

General  description  of  non-publicly  known  private  datasets  obtained  from  third  parties
We  obtain  non-publicly  known  private  datasets  from  third  parties  covering  diverse  domains  and  content  types.

  2.3  Data  crawled  and  scraped  from  online  sources
Were  crawlers  used  by  the  provider  or  on  behalf  of?

Yes

4

Training  Data  Summary   If  yes,  specify  crawler  name(s)/identifier(s):
ClaudeBot

Purposes  of  the  crawler(s):
ClaudeBot  collects  web  content  that  could  potentially  contribute  to  the  model’s  training.
General  description  of  crawler  behaviour:
We  aim  to  minimize  disruption  to  website  owners  and  be  thoughtful  about  how  quickly  ClaudeBot  crawls  domains,    including  by  respecting  crawl-delay  and  disallow  directives  in  robots.txt  files,  where  appropriate,  and  respecting  anti-circumvention  technologies  such  as  paywalls,  password  protection,  and  CAPTCHAs.  More  information  on  our  crawlers  and  how  they  work  can  be  found  at:  https://support.claude.com/en/articles/8896518-does-anthropic-crawl-data-from-the-web-and-how-can-site-owners-block-the-crawler.
Period  of  data  collection:  March  2024  -  July  2026
Comprehensive  description  of  the  type  of  content  and  online  sources  crawled:
The  crawlers  may  be  exposed  to  a  wide  variety  of  content  and  online  sources,  including  most  forms  of  publicly  available  online  data.
Type  of  modality  covered:  Text,  Image

Summary  of  the  most  relevant  domain  names  crawled:
The  portion  of  the  model's  training  corpus  derived  from  data  crawled  and  scraped  from  online  sources  includes  technical  documentation,  open-source  software,  predominantly  text-based  reference  sites,  document  sharing  sites,  and  math  sites.  Top-level  domains  such  as  .com,  .org,  and  .net  are  included  alongside  sites  from  a  range  of  different  countries.
  2.4  User  data   Was  data  from  user  interactions  with  the  AI  model  (e.g.  user  input  and  prompts)  used  to  train  the  model?
 Yes

Was  data  collected  from  user  interactions  with  the  provider’s  other  services  or  products  used  to  train  the  model?
 Yes
If  yes,  provide  a  general  description  of  the  provider’s  services  or  products  that  were  used  to  collect  the  user  data:
 To  the  extent  permitted  by  Anthropic’s  terms  of  service,  privacy  policy,  and  other  contracts,  and  in  line  with  applicable  law,  if  a
5

Training  Data  Summary   user  explicitly  reports  feedback  or  bugs  to  us  (e.g.,  via  thumbs  and  feedback  buttons)  or  otherwise  chooses  to  allow  us  to  use  their  data,  then  chats  and  coding  session  data  may  be  used  in  model  training.  More  information  is  available  at  https://privacy.claude.com/en/articles/7996868-is-my-data-used-for-model-training    We  may  also  incorporate  data  derived  from  Anthropic  employees’  use  of  internal-only  model  versions.
Type  of  modality  covered:  Text

2.5  Synthetic  data   Was  synthetic  AI-generated  data  created  by  the  provider  or  on  their  behalf  to  train  the  model?

 Yes
If  yes,  modality  of  the  synthetic  data:

Text,  Image
If  yes,  specify  the  general-purpose  AI  model(s)  used  to  generate  the  synthetic  data  if  available  on  the  market:

 Synthetic  data  was  provided  by  speech  to  text  models,  large  language  models  (LLM),  and  vision-language  models  (VLM).
 Information  about  other  AI  models,  including  provider’s  own  AI  model(s)  not  available  on  the  market,  used  to  generate  synthetic  data  to  train  the  model  to  which  this  Summary  applies:

 Some  synthetic  data  used  in  training  was  generated  by  Anthropic  models  not  available  on  the  market.

2.6  Other  sources  of  data   Have  data  sources  other  than  those  described  in  Sections  2.1  to  2.5  been  used  to  train  the  model?
Yes

6

Training  Data  Summary   If  yes,  provide  a  narrative  description  of  these  data  sources  and  the  data:
 A  portion  of  the  data  corpus  comes  from  acquired  physical  texts.

1.  Data  processing  aspects
3.  Data  processing  aspects
3.1.

Respect

of

reservation

of

rights

from

text

and

data

mining

exception

or

limitation

  Are  you  a  Signatory  to  the  Code  of  Practice  for  general-purpose  AI  models  that  includes  commitments  to  respect  reservations  of  rights  from  the  TDM  exception  or  limitation?

Yes

Describe  the  measures  implemented  before  model  training  to  respect  reservations  of  rights  from  the  TDM  exception  or  limitation  before  and  during  data  collection,  including  the  opt-out  protocols  and  solutions  honoured  by  the  provider  or,  as  applicable,  by  third  parties  from  which  datasets  have  been  obtained:
 ClaudeBot  respects  crawl-delay  and  disallow  directives  in  robots.txt  files  where  appropriate,  as  well  as  anti-circumvention  technologies  such  as  paywalls,  password  protection,  and  CAPTCHAs.  More  information  on  our  crawlers  and  how  they  work  can  be  found  at:  https://support.claude.com/en/articles/8896518-does-anthropic-crawl-data-from-the-web-and-how-can-site-owners-block-the-crawler.

3.2  Removal  of  illegal  content
General  description  of  measures  taken:
We  take  a  number  of  protective  measures  to  remove  illegal  content  from  the  training  corpus  such  as  active  filtering,  scoring,  moderation,  and  blocking.

3.3.  Other  information  (optional)  Other  relevant  information  about  data  processing  (optional):
Not  applicable

7

Training  Data  Summary
8

===== Claude Sonnet 5 Training Data Summary .pdf =====
Public  Summary  of  Training
Content

Claude

Sonnet

5

Training  Data  Summary
 Version  of  the  Summary:   Version  #1   Last  update:
July

24,

2026
 General  information\

1.  General  information
1.1.

Provider

identification

Provider  name  and  contact  details:
Anthropic  Ireland,  Limited  6th  Floor  South  Bank  House,  Barrow  Street,  Dublin  4,  Dublin  Ireland

Authorised  representative  name  and  contact  details:  Not  applicable

1.2.  Model  identification
Versioned  model  name(s):
 Claude  Sonnet  5  Model  Card:  www.anthropic.com/system-cards
Model  dependencies:  Not  applicable
  Date  of  placement  of  the  model  on  the  Union  market:  June  30,  2026
1.3  Modalities,  overall  training  data  size  and  other  characteristics  Modality  Select  the  modalities  present  in  the  training  data,  to  the  extent  that  they  are
identifiable

Training  data  size  For  each  selected  modality,  select  the  range  within  which  the  estimated  total  training  data  size  for  that  modality  falls.  Dynamic  datasets  may  be  excluded  from  the  estimation.
Types  of  content  For  each  selected  modality,  provide  a  general  description  of  the  type  of  content  that  has  been  included  in  the
training

data.

Text

☐

Less

than

1

billion

tokens
 ☐  1billion  to  10  trillions  tokens  X  More  than  10  trillions  tokens
The  training  corpus  for  the  model  includes  an  array  of  text  types,  including  short  and  long-form  texts,  software  code,  synthetic  text,  prose  in  a  variety  of  languages,  mathematical  data,  and  prompts  and  preference  data  used  during  reinforcement  learning.
2

Training  Data  Summary
Image
☐

Less

than

1

million

images

☐

1Million

to1

billion

images
 X  More  than  1  billion  images
The  training  corpus  for  the  model  includes  an  array  of  image  types,  including  photographs,  interleaved  text  and  images  from  websites,  computer  graphics,  and  prompts  and  preference  data  used  during  reinforcement  learning.   Video  Not  applicable  Audio  Not  applicable  Other  Not  applicable
 Latest  date  of  data  acquisition/collection  for  model  training:
A  number  of  different  datasets,  with  varying  publication  and  cut-off  dates,  are  included  in  the  training  corpus,  with  some  data  being  acquired/collected  up  to  May  2026.
Description  of  the  linguistic  characteristics  of  the  overall  training  data:
Training  sources  deliberately  include  a  diverse  range  of  global  languages,  both  European  and  non-European,  including  those  with  relatively  high  numbers  of  speakers  ( e.g.,  English,  Chinese,  French,  Spanish)  as  well  as  comparably  lower  volumes  ( e.g.,  Basque,  Breton,  Korean).
2.  List  of  data  sources
2.  List  of  data  sources
2.1.  Publicly  available  datasets
Have  you  used  publicly  available  datasets  to  train  the  model?    Yes
If  yes,  specify  the  modality(ies)  of  the  content  covered  by  the  datasets  concerned:
Text,  Image
List  of  large  publicly  available  datasets:
The  training  corpus  is  derived  from  several  publicly  accessible  repositories,  notably  including  Common  Crawl,  a  repository  of  web  crawl  data,  as  well  as  specialized  datasets  available  through  platforms  like  GitHub  and  HuggingFace.
General  description  of  other  publicly  available  datasets  not  listed  above:
Data  from  other  publicly  available  datasets  is  included  in  the  training  corpus,  including  mathematical  data  and  some  image  and  caption  datasets.

2.2  Private  non-publicly  available  datasets  obtained  from  third
parties

3

Training  Data  Summary
2.2.1.  Datasets  commercially  licensed  by  rightsholders  or  their  representatives  Have  you  concluded  transactional  commercial  licensing  agreement(s)  with  rightsholder(s)  or  with  their  representatives?
Yes
If  yes,  specify  the  modality(ies)  of  the  content  covered  by  the  datasets  concerned:   Text,  Image
2.2.2.  Private  datasets  obtained  from  other  third  parties  Have  you  obtained  private  datasets  from  third  parties  that  are  not  licensed  as  described  in  Section  2.2.1,  such  as  data  obtained  from  providers  of  private  databases,  or  data  intermediaries?
Yes
If  yes,  specify  the  modality(ies)  of  the  content  covered  by  the  datasets  concerned:
Text,  Image

If  publicly  known,  list  private  datasets  obtained  from  other  third  parties:   Not  applicable

General  description  of  non-publicly  known  private  datasets  obtained  from  third  parties
We  obtain  non-publicly  known  private  datasets  from  third  parties  covering  diverse  domains  and  content  types.

2.3  Data  crawled  and  scraped  from  online  sources
Were  crawlers  used  by  the  provider  or  on  behalf  of?

Yes

If  yes,  specify  crawler  name(s)/identifier(s):
ClaudeBot

Purposes  of  the  crawler(s):
ClaudeBot  collects  web  content  that  could  potentially  contribute  to  the  model’s  training.
General  description  of  crawler  behaviour:
We  aim  to  minimize  disruption  to  website  owners  and  be  thoughtful  about  how  quickly  ClaudeBot  crawls  domains,    including  by  respecting  crawl-delay  and  disallow  directives  in  robots.txt  files,  where  appropriate,  and  respecting  anti-circumvention  technologies  such  as  paywalls,  password  protection,  and  CAPTCHAs.  More  information  on  our  crawlers  and  how  they  work  can  be  found  at:  https://support.claude.com/en/articles/8896518-does-anthropic-crawl-data-from-the-web-and-how-can-site-owners-block-the-crawler.
4

Training  Data  Summary   Period  of  data  collection:  March  2024  -  May  2026
Comprehensive  description  of  the  type  of  content  and  online  sources  crawled:
The  crawlers  may  be  exposed  to  a  wide  variety  of  content  and  online  sources,  including  most  forms  of  publicly  available  online  data.
Type  of  modality  covered:  Text,  Image

Summary  of  the  most  relevant  domain  names  crawled:
The  portion  of  the  model's  training  corpus  derived  from  data  crawled  and  scraped  from  online  sources  includes  technical  documentation,  open-source  software,  predominantly  text-based  reference  sites,  document  sharing  sites,  and  math  sites.  Top-level  domains  such  as  .com,  .org,  and  .net  are  included  alongside  sites  from  a  range  of  different  countries.
2.4  User  data   Was  data  from  user  interactions  with  the  AI  model  (e.g.  user  input  and  prompts)  used  to  train  the  model?
 Yes

Was  data  collected  from  user  interactions  with  the  provider’s  other  services  or  products  used  to  train  the  model?
 Yes
If  yes,  provide  a  general  description  of  the  provider’s  services  or  products  that  were  used  to  collect  the  user  data:
 To  the  extent  permitted  by  Anthropic’s  terms  of  service,  privacy  policy,  and  other  contracts,  and  in  line  with  applicable  law,  if  a  user  explicitly  reports  feedback  or  bugs  to  us  (e.g.,  via  thumbs  and  feedback  buttons)  or  otherwise  chooses  to  allow  us  to  use  their  data,  then  chats  and  coding  session  data  may  be  used  in  model  training.  More  information  is  available  at  https://privacy.claude.com/en/articles/7996868-is-my-data-used-for-model-training    We  may  also  incorporate  data  derived  from  Anthropic  employees’  use  of  internal-only  model  versions.
Type  of  modality  covered:  Text

2.5  Synthetic  data   Was  synthetic  AI-generated  data  created  by  the

 Yes
5

Training  Data  Summary   provider  or  on  their  behalf  to  train  the  model?
If  yes,  modality  of  the  synthetic  data:

Text,  Image
If  yes,  specify  the  general-purpose  AI  model(s)  used  to  generate  the  synthetic  data  if  available  on  the  market:

 Synthetic  data  was  provided  by  speech  to  text  models,  large   language  models  (LLM),  and  vision-language  models  (VLM).
 Information  about  other  AI  models,  including  provider’s  own  AI  model(s)  not  available  on  the  market,  used  to  generate  synthetic  data  to  train  the  model  to  which  this  Summary  applies:

 Some  synthetic  data  used  in  training  was  generated  by  Anthropic  models  not  available  on  the  market.

2.6  Other  sources  of  data   Have  data  sources  other  than  those  described  in  Sections  2.1  to  2.5  been  used  to  train  the  model?
Yes

If  yes,  provide  a  narrative  description  of  these  data  sources  and  the  data:
 A  portion  of  the  data  corpus  comes  from  acquired  physical  texts.

1.  Data  processing  aspects
3.  Data  processing  aspects
3.1.

Respect

of

reservation

of

rights

from

text

and

data

mining

exception

or

limitation

  Are  you  a  Signatory  to  the  Code  of  Practice  for  general-purpose  AI  models  that  includes  commitments  to  respect  reservations  of  rights  from  the  TDM  exception  or  limitation?

Yes

6

Training  Data  Summary
Describe  the  measures  implemented  before  model  training  to  respect  reservations  of  rights  from  the  TDM  exception  or  limitation  before  and  during  data  collection,  including  the  opt-out  protocols  and  solutions  honoured  by  the  provider  or,  as  applicable,  by  third  parties  from  which  datasets  have  been  obtained:
 ClaudeBot  respects  crawl-delay  and  disallow  directives  in  robots.txt  files  where  appropriate,  as  well  as  anti-circumvention  technologies  such  as  paywalls,  password  protection,  and  CAPTCHAs.  More  information  on  our  crawlers  and  how  they  work  can  be  found  at:  https://support.claude.com/en/articles/8896518-does-anthropic-crawl-data-from-the-web-and-how-can-site-owners-block-the-crawler.

3.2  Removal  of  illegal  content
General  description  of  measures  taken:
We  take  a  number  of  protective  measures  to  remove  illegal  content  from  the  training  corpus  such  as  active  filtering,  scoring,  moderation,  and  blocking.

3.3.  Other  information  (optional)  Other  relevant  information  about  data  processing  (optional):
Not  applicable

7

===== _Claude Mythos Preview Training Data Summary .pdf =====
Public  Summary  of  Training
Content

Claude

Mythos

Preview

Training  Data  Summary

  Version  of  the  Summary:   Version  #1
Last  update:

July

24,

2026
 General  information

1.  General  information
1.1.

Provider

identification

Provider  name  and  contact  details:
Anthropic  Ireland,  Limited  6th  Floor  South  Bank  House,  Barrow  Street,  Dublin  4,  Dublin  Ireland

Authorised  representative  name  and  contact  details:  Not  applicable

1.2.  Model  identification
Versioned  model  name(s):
 Claude  Mythos  Preview  Model  Card:  www.anthropic.com/system-cards
Model  dependencies:  Not  applicable
  Date  of  placement  of  the  model  on  the  Union  market:  June  2,  2026
1.3  Modalities,  overall  training  data  size  and  other  characteristics  Modality  Select  the  modalities  present  in  the  training  data,  to  the  extent  that  they  are
identifiable

Training  data  size  For  each  selected  modality,  select  the  range  within  which  the  estimated  total  training  data  size  for  that  modality  falls.  Dynamic  datasets  may  be  excluded  from  the  estimation.
Types  of  content  For  each  selected  modality,  provide  a  general  description  of  the  type  of  content  that  has  been  included  in  the
training

data.

2

Training  Data  Summary
Text

☐

Less

than

1

billion

tokens
 ☐  1billion  to  10  trillions  tokens  X  More  than  10  trillions  tokens
The  training  corpus  for  the  model  includes  an  array  of  text  types,  including  short  and  long-form  texts,  software  code,  synthetic  text,  prose  in  a  variety  of  languages,  mathematical  data,  and  prompts  and  preference  data  used  during  reinforcement  learning.
Image
☐

Less

than

1

million

images

☐

1Million

to1

billion

images
 X  More  than  1  billion  images
The  training  corpus  for  the  model  includes  an  array  of  image  types,  including  photographs,  interleaved  text  and  images  from  websites,  computer  graphics,  and  prompts  and  preference  data  used  during  reinforcement  learning.   Video  Not  applicable  Audio  Not  applicable  Other  Not  applicable
  Latest  date  of  data  acquisition/collection  for  model  training:
A  number  of  different  datasets,  with  varying  publication  and  cut-off  dates,  are  included  in  the  training  corpus,  with  some  data  being  acquired/collected  up  to  February  2026.
Description  of  the  linguistic  characteristics  of  the  overall  training  data:
Training  sources  deliberately  include  a  diverse  range  of  global  languages,  both  European  and  non-European,  including  those  with  relatively  high  numbers  of  speakers  ( e.g.,  English,  Chinese,  French,  Spanish)  as  well  as  comparably  lower  volumes  ( e.g.,  Basque,  Breton,  Korean).
2.  List  of  data  sources
2.  List  of  data  sources
2.1.  Publicly  available  datasets
Have  you  used  publicly  available  datasets  to  train  the  model?    Yes
If  yes,  specify  the  modality(ies)  of  the  content  covered  by  the  datasets  concerned:
Text,  Image
List  of  large  publicly  available  datasets:
The  training  corpus  is  derived  from  several  publicly  accessible  repositories,  notably  including  Common  Crawl,  a  repository  of  web  crawl  data,  as  well  as  specialized  datasets  available  through  platforms  like  GitHub  and  HuggingFace.
3

Training  Data  Summary   General  description  of  other  publicly  available  datasets  not  listed  above:
Data  from  other  publicly  available  datasets  is  included  in  the  training  corpus,  including  mathematical  data  and  some  image  and  caption  datasets.

2.2

Private

non-publicly

available

datasets

obtained

from

third

parties
  2.2.1.  Datasets  commercially  licensed  by  rightsholders  or  their  representatives  Have  you  concluded  transactional  commercial  licensing  agreement(s)  with  rightsholder(s)  or  with  their  representatives?
Yes
If  yes,  specify  the  modality(ies)  of  the  content  covered  by  the  datasets  concerned:   Text,  Image

2.2.2.

Private

datasets

obtained

from

other

third

parties
 Have  you  obtained  private  datasets  from  third  parties  that  are  not  licensed  as  described  in  Section  2.2.1,  such  as  data  obtained  from  providers  of  private  databases,  or  data  intermediaries?
Yes
If  yes,  specify  the  modality(ies)  of  the  content  covered  by  the  datasets  concerned:
Text,  Image

If  publicly  known,  list  private  datasets  obtained  from  other  third  parties:   Not  applicable

General  description  of  non-publicly  known  private  datasets  obtained  from  third  parties
We  obtain  non-publicly  known  private  datasets  from  third  parties  covering  diverse  domains  and
content

types.

2.3  Data  crawled  and  scraped  from  online  sources
Were  crawlers  used  by  the  provider  or  on  behalf  of?

Yes

If  yes,  specify  crawler  name(s)/identifier(s):
ClaudeBot

Purposes  of  the  crawler(s):
ClaudeBot  collects  web  content  that  could  potentially  contribute  to  the  model’s  training.
4

Training  Data  Summary
General  description  of  crawler  behaviour:
We  aim  to  minimize  disruption  to  website  owners  and  be  thoughtful  about  how  quickly  ClaudeBot  crawls  domains,    including  by  respecting  crawl-delay  and  disallow  directives  in  robots.txt  files  where  appropriate,  and  respecting  anti-circumvention  technologies  such  as  paywalls,  password  protection,  and  CAPTCHAs.  More  information  on  our  crawlers  and  how  they  work  can  be  found  at:  https://support.claude.com/en/articles/8896518-does-anthropic-crawl-data-from-the-web-and-how-can-site-owners-block-the-crawler.
Period  of  data  collection:  March  2024  -  March  2026
Comprehensive  description  of  the  type  of  content  and  online  sources  crawled:
The  crawlers  may  be  exposed  to  a  wide  variety  of  content  and  online  sources,  including  most  forms  of  publicly  available  online  data.
Type  of  modality  covered:  Text,  Image

Summary  of  the  most  relevant  domain  names  crawled:
The  portion  of  the  model's  training  corpus  derived  from  data  crawled  and  scraped  from  online  sources  includes  technical  documentation,  open-source  software,  predominantly  text-based  reference  sites,  document  sharing  sites,  and  math  sites.  Top-level  domains  such  as  .com,  .org,  and  .net  are  included  alongside  sites  from  a  range  of  different  countries.
2.4  User  data   Was  data  from  user  interactions  with  the  AI  model  (e.g.  user  input  and  prompts)  used  to  train  the  model?
 Yes

Was  data  collected  from  user  interactions  with  the  provider’s  other  services  or  products  used  to  train  the  model?
 Yes
If  yes,  provide  a  general  description  of  the  provider’s  services  or  products  that  were  used  to  collect  the  user  data:
 To  the  extent  permitted  by  Anthropic’s  terms  of  service,  privacy  policy,  and  other  contracts,  and  in  line  with  applicable  law,  if  a  user  explicitly  reports  feedback  or  bugs  to  us  (e.g.,  via  thumbs  and  feedback  buttons)  or  otherwise  chooses  to  allow  us  to  use  their  data,  then  chats  and  coding  session  data  may  be  used  in  model  training.  More  information  is  available  at  https://privacy.claude.com/en/articles/7996868-is-my-data-used-for-model-training    We  may  also  incorporate  data  derived  from
5

Training  Data  Summary   Anthropic  employees’  use  of  internal-only  model  versions.
Type  of  modality  covered:  Text

2.5  Synthetic  data   Was  synthetic  AI-generated  data  created  by  the  provider  or  on  their  behalf  to  train  the  model?

 Yes
If  yes,  modality  of  the  synthetic  data:

Text,  Image
If  yes,  specify  the  general-purpose  AI  model(s)  used  to  generate  the  synthetic  data  if  available  on  the  market:

 Synthetic  data  was  provided  by  speech  to  text  models,  large  language  models  (LLM),  and  vision-language  models  (VLM).
 Information  about  other  AI  models,  including  provider’s  own  AI  model(s)  not  available  on  the  market,  used  to  generate  synthetic  data  to  train  the  model  to  which  this  Summary  applies:

 Some  synthetic  data  used  in  training  was  generated  by  Anthropic  models  not  available  on  the  market.

2.6  Other  sources  of  data   Have  data  sources  other  than  those  described  in  Sections  2.1  to  2.5  been  used  to  train  the  model?
Yes

If  yes,  provide  a  narrative  description  of  these  data  sources  and  the  data:
 A  portion  of  the  data  corpus  comes  from  acquired  physical  texts.

6

Training  Data  Summary
1.
3.  Data  processing  aspects
3.1.

Respect

of

reservation

of

rights

from

text

and

data

mining

exception

or

limitation

  Are  you  a  Signatory  to  the  Code  of  Practice  for  general-purpose  AI  models  that  includes  commitments  to  respect  reservations  of  rights  from  the  TDM  exception  or  limitation?

Yes

Describe  the  measures  implemented  before  model  training  to  respect  reservations  of  rights  from  the  TDM  exception  or  limitation  before  and  during  data  collection,  including  the  opt-out  protocols  and  solutions  honoured  by  the  provider  or,  as  applicable,  by  third  parties  from  which  datasets  have  been  obtained:
 ClaudeBot  respects  crawl-delay  and  disallow  directives  in  robots.txt  files  where  appropriate,  as  well  as  anti-circumvention  technologies  such  as  paywalls,  password  protection,  and  CAPTCHAs.  More  information  on  our  crawlers  and  how  they  work  can  be  found  at:  https://support.claude.com/en/articles/8896518-does-anthropic-crawl-data-from-the-web-and-how-can-site-owners-block-the-crawler.

3.2  Removal  of  illegal  content
General  description  of  measures  taken:
We  take  a  number  of  protective  measures  to  remove  illegal  content  from  the  training  corpus  such  as  active  filtering,  scoring,  moderation,  and  blocking.

3.3.  Other  information  (optional)  Other  relevant  information  about  data  processing  (optional):
Not  applicable.

7