SciDraw-6K: A Multilingual Scientific Illustration Dataset Generated by Google Gemini

Published: 20 April 2026| Version 1 | DOI: 10.17632/m8z2jyr7z5.1
Contributor:
Davie Chen

Description

SciDraw-6K is a curated dataset of 6,291 scientific illustrations synthesized by Google Gemini image-generation models (primarily gemini-3-pro-image-preview and gemini-2.5-flash-image), each paired with aligned prompts in eleven languages: English, Simplified Chinese, Traditional Chinese, Japanese, Korean, German, French, Spanish, Brazilian Portuguese, Italian, and Russian. Images span eight broad scientific categories — biomedical (44.9%), materials (13.4%), AI systems (11.2%), chemistry (9.7%), environment (9.2%), electronics (3.0%), physics (2.2%), and a residual "other" bucket (6.3%) covering long-tail disciplines such as robotics, mathematics, economics, civil engineering, and geosciences. Unlike general-purpose text-to-image corpora (LAION-5B, JourneyDB, DiffusionDB) which are dominated by photorealistic and artistic content, SciDraw-6K is purpose-built for the scientific-illustration genre: schematic diagrams, mechanism figures, table-of-contents graphical abstracts, and conceptual posters. The dataset is constructed via a domain-specific prompt taxonomy, Gemini image generation, LLM-based translation, and lightweight quality control. Each row of the metadata contains: a stable image ID, the public image URL, the file extension, the category label, the eleven multilingual prompts, the Gemini model identifier, the generation type, the creation timestamp, and the SHA-256 hash of the downloaded image bytes. Intended uses include: multilingual text-to-image research, domain-adapted diffusion fine-tuning, prompt-engineering studies for scientific visualization, retrieval-augmented generation for scientific figure synthesis, and benchmarking frontier image-generation models on the specialized visual grammar of science. This Mendeley Data record archives the dataset metadata. The full image payload (~19 GB) is hosted on Hugging Face (https://huggingface.co/datasets/SciDrawAI/SciDraw-6K) and archived on Zenodo (https://doi.org/10.5281/zenodo.19642870) and Harvard Dataverse (https://doi.org/10.7910/DVN/L02REW). Construction scripts are available at https://github.com/SciDrawAI/scidraw-6k. Service powered by this dataset: https://sci-draw.com.

Files

Steps to reproduce

1. Collect a user-supplied research topic through the sci-draw.com authoring interface. 2. Compose a concrete prompt by filling the topic into one of the scientific-illustration prompt templates (categorized across biomedical, chemistry, materials, electronics, environment, AI systems, physics, and other). 3. Send the prompt to the Google Gemini image-generation API (gemini-3-pro-image-preview, gemini-2.5-flash-image, or gemini-3.1-flash-image-preview) under default safety and resolution settings. 4. Translate the original English prompt into 10 additional languages via an LLM-based translation pipeline. 5. Apply lightweight quality control: filter out generation failures, off-topic outputs, and policy violations. 6. Strip all user identifiers and session metadata. 7. Compute SHA-256 hash of the downloaded image bytes for integrity verification. 8. Export metadata to JSONL and Parquet formats; organize images by category. Full construction scripts are available at: https://github.com/SciDrawAI/scidraw-6k

Categories

Image Artifact

Licence