HiPuSi-MT44K: Multi-Engine Machine Translation Outputs for Hindi, Punjabi and Devanagari-Script Sindhi in Six Translation Directions

Published: 12 August 2026| Version 1 | DOI: 10.17632/drb9ftgmhb.1
Contributors:
Palak Arora, Bharti Nathani, Nisheeth Joshi

Description

HiPuSi-MT44K is a multi-engine machine translation output dataset for Hindi, Punjabi, and Devanagari-script Sindhi. It contains 3,000 source sentences and 44,000 machine-generated translation outputs across six translation directions. The resource includes outputs from web-based MT services, conversational large language models, pretrained multilingual neural machine translation systems, and in-house EBMT, SMT, and NMT systems. The source collection is general-domain material compiled during a MeitY, Government of India-sponsored project and manually vetted for grammatical structure and spelling by five native-speaker language experts holding Master's-level or higher qualifications. The dataset is intended for comparative MT evaluation, quality estimation, error analysis, system agreement studies, and human evaluation research.

Files

Steps to reproduce

1. Collect 1,000 Hindi, 1,000 Punjabi, and 1,000 Sindhi source sentences from the general-domain source collection. The original collection was compiled from multiple online repositories, including material from Wikipedia and the Prime Minister of India’s *Mann Ki Baat* programme. 2. Manually vet the source sentences before translation. In the original study, five native-speaker language experts with Master’s-level or higher qualifications checked the grammatical structure and spelling of the source text. 3. Translate each source-language collection into the other two languages, producing six translation directions: Hindi→Punjabi, Punjabi→Hindi, Hindi→Sindhi, Sindhi→Hindi, Punjabi→Sindhi, and Sindhi→Punjabi. 4. For Google Translate, Microsoft/Bing Translator, Devnagri Translator, Anuvaad/AI4Bharat, and HimangY, submit the source sentences manually through the corresponding online translation interfaces and record the returned translations. These systems were accessed between July 2024 and March 2025. 5. For ChatGPT and Gemini, submit the source sentences through their conversational web interfaces using the generic prompt: “Translate in xx language”, where “xx” is replaced by the required target language. Record the generated translations manually. These systems were also accessed between July 2024 and March 2025. 6. Generate NLLB outputs locally using the `facebook/nllb-200-3.3B` checkpoint (3.3 billion parameters) and the example inference code provided with its Hugging Face repository. 7. Generate IndicTrans2 outputs locally using the `ai4bharat/indictrans2-indic-indic-1B` checkpoint (1 billion parameters) and the example inference code supplied with its Hugging Face repository. 8. Generate the in-house MT outputs using: (a) the Sataanuvaadak/EBMT system, which stores translation examples as phrases and chunks in a vector database; (b) the statistical MT system trained using Moses RELEASE-4.0; and (c) the Hemant NMT system trained using OpenNMT-py. 9. Preserve the original system outputs without lexical or grammatical correction. Convert the direction-wise spreadsheets into long-format records and apply Unicode NFC normalization with leading and trailing whitespace removal. 10. Generate the dataset quality-control fields, including source-duplicate flags, exact-source-copy flags, abnormal-length flags, and cross-script mixing flags. 11. Run the supplied `scripts/validate_dataset.py` script to verify the expected 44,000 records, direction-level counts, non-empty outputs, unique record identifiers, and source-record counts. Verify file integrity using the provided `SHA256SUMS.txt`. Because the web-based translation services and conversational LLM interfaces are continuously updated, rerunning them at a later date may not reproduce the exact translations contained in this dataset. The released raw files therefore represent the outputs observed during the July 2024–March 2025 collection period.

Categories

Computer Science, Artificial Intelligence, Natural Language Processing

Funders

  • Anusandhan National Research Foundation
    Grant ID: SPG/2021/003306

Licence