Ponys.ai Audited Multilingual AI Character Catalog v1.0.0
Description
Versioned multilingual research catalog of audited Ponys.ai character pages for reproducible AI companion discovery and localization studies. The release contains 104 records with stable character slugs, canonical URLs, language and market labels, content categories, and validation metadata. Coverage emphasizes Japanese, Korean, Spanish for Latin America, Portuguese for Brazil, Simplified Chinese, Traditional Chinese, and English. Included formats are CSV, JSON, MLCommons Croissant metadata, a data dictionary, checksums, and CC BY 4.0 licensing. Intended uses include multilingual retrieval evaluation, localization QA, character taxonomy research, benchmark construction, and feed or widget integration. Source records were validated against public Ponys.ai pages before release.
Files
Steps to reproduce
1. Verify file integrity with checksums.json. 2. Load characters.csv in a UTF-8 aware tabular tool or characters.json in an application. 3. Join records by slug. 4. Validate each canonical_url with HTTP 200 and self-referencing canonical checks. 5. Group by language and market for JP, KR, LATAM and BR evaluation. 6. Use croissant.json for machine-readable dataset discovery. 7. Cite the versioned DOI and preserve the CC BY 4.0 attribution.