GPAI Ledger › GPAI Training Transparency tracker (AI Accountability Lab (AIAL)) › Capture 12 Sep 2026
ByteDance_Seedeam_5_0_2026_09_08 — capture 20260912T062129Z
Filed under AI Accountability Lab (AIAL) — GPAI Training Transparency tracker, the source this project captured it from; the document itself is the filing of the model named above.
| Provider | provider not identified by this project |
|---|---|
| Target | AIAL archived copy — https://raw.githubusercontent.com/AIAccountabilityLab/gpai-training-transparency/c878ebbe5a3d7585857ed203f59f34551d7e72f7/public/archive/ByteDance_Seedeam_5_0_2026_09_08.pdf |
| Fetched (UTC) | 2026-09-12T06:21:29Z |
| Upstream commit | 8 Sep 2026 — c878ebbe5a3d (when this state began to stand in the upstream repository; this archive fetched it at the time above, not then) |
| Stored file | bbfe07316d3bfeb78f67e43c8280e887102aa2cb4926db18e1a3c9fa25fa1e7e.pdf (251,620 bytes) |
| SHA-256 | bbfe07316d3bfeb78f67e43c8280e887102aa2cb4926db18e1a3c9fa25fa1e7e |
| OpenTimestamps proof | bbfe07316d3bfeb78f67e43c8280e887102aa2cb4926db18e1a3c9fa25fa1e7e.pdf.20260912T062129Z.ots (calendar-attested; anchored in bitcoin over time) |
| Wayback | not saved |
| Prior capture of this target | — first capture of this target |
Verify: sha256sum bbfe07316d3bfeb78f67e43c8280e887102aa2cb4926db18e1a3c9fa25fa1e7e.pdf must equal the hash above (the filename IS the expected hash); ots verify bbfe07316d3bfeb78f67e43c8280e887102aa2cb4926db18e1a3c9fa25fa1e7e.pdf.20260912T062129Z.ots -f bbfe07316d3bfeb78f67e43c8280e887102aa2cb4926db18e1a3c9fa25fa1e7e.pdf (opentimestamps.org) proves the bytes existed no later than the attestation time — an upper bound on the capture time; the fetch time above is the archive's own record (a freshly captured proof reports 'pending' here: the calendars anchor within hours, but this archive only upgrades the stored proof to its anchor on a later run, so expect a day or two). ots verify needs a local Bitcoin Core node (a pruned one is fine); without one, ots info on the proof prints the attesting block height and merkle path to check on any block explorer.
Extracted text
Machine-extracted text (layout may be lost; the authoritative content is the stored file above).
PublicSummaryofTrainingDataContent forSeedream5.0Pro VersionoftheSummary: 1.0 Lastupdate: 31July2026 1. Generalinformation 1.1 Provideridentification Providernameandcontact details: ByteDanceNexusAIPte.Ltd(ByteDance) Authorisedrepresentative nameandcontactdetails: MikrosInformationTechnologyIrelandLimited 1.2 Modelidentification Versionedmodelname(s): TheSeedream5.0Profamilyofmodels Modeldependencies: TheSeedream5.0ProfamilyincludesSeeDream5.0ProandSeedream5.0 Lite.Bothmodelsarebasedonthesamearchitecture. Dateofplacementofthe modelontheUnion market: July2026 1.3 Modalities,overalltrainingdatasizeandothercharacteristics Modality Trainingdatasize Typesofcontent Text Morethan10trilliontokens Seedream5.0Pro'strainingdataset comprisespairedimage-textdatasets andinterweavedimage-textdata, includingimageswithcorresponding textualdescriptionsanddocuments withinterweavedimagesandtext. Image Morethan1billionimages Seedream5.0Pro'strainingdataset comprisespairedimage-textdatasets andinterweavedimage-textdata, includingimageswithcorresponding textualdescriptionsanddocuments withinterweavedimagesandtext. Audio N/A N/A Video N/A N/A Other N/A N/A Latestdateofdata acquisition/collectionfor modeltraining: June2026 Descriptionofthelinguistic characteristicsoftheoverall trainingdata: Nospecificgeographicregionwasintentionallyexcludedfromthe datacollectionprocess. Otherrelevantcharacteristics oftheoveralltrainingdata: Seedream5.0Pro'strainingdatasetcomprisespairedimage-text datasetsandinterweavedimage-textdata,includingimageswith correspondingtextualdescriptionsanddocumentswithinterweaved imagesandtext. Formodeltrainingpurposes,ByteDanceusesalarge-scalemixtureof publiclyavailabledata,licenseddata,andsyntheticdata,designedto ensurebroadcoverageofdomainsandcontexts. Thedatasethasbeencuratedwiththeobjectiveofmaximizing representativenessandrobustness,whileapplyingfiltering mechanismstoremovelow-qualityorharmfulcontentwhere appropriate. Additionalcomments (optional): N/A 2. ListofDataSources 2.1 Publiclyavailabledatasets Haveyouusedpubliclyavailable datasetstotrainthemodel? Yes Ifyes,specifythemodality(ies) ofthecontentcoveredbythe datasetsconcerned: Image Listoflargepubliclyavailable datasets: Thetrainingdatacorpuscomprisesamixofpubliclyavailableand licenseddata. Thismayincludeimagedatasetsmadeavailablebythirdparties throughpublicrepositoriesandonlineplatforms.Asdescribed above,suchdatasetsarecuratedandpre-processedwiththe objectiveofmaximizingrepresentativenessandrobustness,while applyingfilteringmechanismstoremovelow-qualityorharmful contentwhereappropriate. Generaldescriptionofother publiclyavailabledatasetsnot listedabove: Seethedescriptionoflinguisticcharacteristicsandotherrelevant characteristicsofoveralltrainingdatainSection1. Additionalcomments(optional): N/A 2.2 Privatenon-publiclyavailabledatasetsobtainedfromthird parties 2.2.1 Datasetscommerciallylicensedbyrightsholdersortheir representatives Haveyouconcludedtransactionalcommerciallicensing agreement(s)withrightsholder(s)orwiththeir representatives? Yes Image Ifyes,specifythemodality(ies)ofthecontentcoveredby thedatasetsconcerned: 2.2.2 Privatedatasetsobtainedfromotherthirdparties Haveyouobtainedprivatedatasetsfrom thirdpartiesthatarenotlicensedas describedinSection2.2.1,suchasdata obtainedfromprovidersofprivatedatabases, ordataintermediaries? Yes Ifyes,specifythemodality(ies)ofthecontent coveredbythedatasetsconcerned: Image Ifpubliclyknown,listprivatedatasets obtainedfromotherthirdparties: N/A Generaldescriptionofnon-publiclyknown privatedatasetsobtainedfromthirdparties Wehaveobtaineddatasetsfromthird-partiesona licensedbasis.Suchdatasetsprimarilycomprise contentacrosstextandimagemodalities. Additionalcomments(optional): N/A 2.3 Datacrawledandscrapedfromonlinesources Werecrawlersusedbytheprovideror onbehalfof? Yes Ifyes,specifycrawler name(s)/identifier(s): Bytespider Purposesofthecrawler(s): InthecontextofByteDance'sAImodeldevelopment, Bytespiderisusedtocollectpubliclyaccessibleimage,video andaudiodatatosupporttraining,testing,andvalidationof theAImodel. Generaldescriptionofcrawler behaviour: ByteDance'scrawlerisconfiguredtoavoidoverloading websitesandtooperateinamannerconsistentwithethical webscrapingpractices.Ourcrawlerisdesignedtorespect robots.txtinstructions,andnottocircumventcontrol measuressuchaspaywallsortoaccesspassword-protected content. Periodofdatacollection: UptoJune2026 Comprehensivedescriptionofthetype ofcontentandonlinesourcescrawled: Asabove,wemayuseourcrawlertocollectpublicly accessibledata.Thecrawledcontentprimarilyconsistsof videos,imagesandaudio. Typeofmodalitycovered: Image Summaryofthemostrelevantdomain namescrawled: Themostrelevantdomainscrawledincludesresources spanningnaturalscenes,objects,peopleanddiagrams. Additionalcomments(optional): N/A 2.4 Userdata WasdatafromuserinteractionswiththeAImodel (e.g.userinputandprompts)usedtotrainthe model? No Wasdatacollectedfromuserinteractionswiththe provider’sotherservicesorproductsusedtotrain themodel? No Ifyes,provideageneraldescriptionofthe provider’sservicesorproductsthatwereusedto collecttheuserdata: N/A Typeofmodalitycovered: N/A Additionalcomments(optional): N/A 2.5 Syntheticdata WassyntheticAI-generateddatacreatedbythe providerorontheirbehalftotrainthemodel? Yes Ifyes,modalityofthesyntheticdata: Textandimage Ifyes,specifythegeneral-purposeAImodel(s)used togeneratethesyntheticdataifavailableonthe market: Syntheticdatawasgeneratedusingarangeof general-purposeAImodels,includingimage generationmodels,visionlanguagemodelsand largelanguagemodels. InformationaboutotherAImodels,including provider’sownAImodel(s)notavailableonthe market,usedtogeneratesyntheticdatatotrainthe modeltowhichthisSummaryapplies: Wemayusefine-tunedversionsofinternal modelstogeneratesyntheticdata. Additionalcomments(optional): N/A 2.6 Othersourcesofdata Havedatasourcesotherthanthosedescribedin Sections2.1to2.5beenusedtotrainthemodel? No Ifyes,provideanarrativedescriptionofthesedata sourcesandthedata: N/A Additionalcomments(optional): N/A 3. Dataprocessingaspects 3.1 Respectofreservationofrightsfromtextanddatamining exceptionorlimitation AreyouaSignatorytotheCodeofPracticeforgeneral- purposeAImodelsthatincludescommitmentstorespect reservationsofrightsfromtheTDMexceptionorlimitation? No Describethemeasuresimplementedbeforemodeltraining torespectreservationsofrightsfromtheTDMexceptionor limitationbeforeandduringdatacollection,includingthe opt-outprotocolsandsolutionshonouredbytheprovider or,asapplicable,bythirdpartiesfromwhichdatasetshave beenobtained: ByteDance'scrawlerisconfiguredtoavoid overloadingwebsitesandtooperateina mannerconsistentwithethicalweb scrapingpractices.Ourcrawleris designedtorespectrobots.txt instructions,andnottocircumvent controlmeasuressuchaspaywallsorto accesspassword-protectedcontent. Additionalcomments(optional): N/A 3.2 Removalofillegalcontent Generaldescriptionofmeasurestaken: Asmentioned,weappliedpreprocessingandfiltering methods,includingfilteringmodels,toexcludeunsafeor harmfulcontent. 3.3 Otherinformation(optional) Otherrelevantinformationaboutdata processing(optional): N/A