Myanmar Written-Spoken Text Pairs
A parallel text dataset pairing formal written Myanmar sentences with their natural spoken alternatives for style transfer and NLP tasks.
csvLicense: CC-BY-NC-SA-4.0A centralized, open-access catalog of structured datasets covering Myanmar NLP, geospatial mapping, humanitarian metrics, and technical benchmarks. Maintained by the community for researchers and machine learning practitioners.
A parallel text dataset pairing formal written Myanmar sentences with their natural spoken alternatives for style transfer and NLP tasks.
csvLicense: CC-BY-NC-SA-4.0An aligned English-Burmese parallel sentence corpus extracted from Hugging Face Course video subtitles to support machine translation research.
csvLicense: CC-BY-4.0A high-quality, human-verified parallel corpus covering general domain texts, everyday conversations, and literature for Myanmar-English machine translation.
parquetLicense: CC-BY-NC-SA-4.0A curated text corpus of grammatically sound Burmese sentences collected from news and educational websites for general NLP tasks.
jsonl, textLicense: CC-BY-4.0A diverse synthetic image-to-text dataset featuring Burmese characters rendered in various visual styles, colors, and backgrounds for OCR training.
parquet, image, textLicense: CC-BY-4.0A labeled corpus designed for grammatical error correction and disambiguation of the frequently confused Burmese particles 'လဲ' and 'လည်း'.
csvLicense: CC-BY-4.0An open-source text corpus containing general-domain sentences written in natural Burmese to support language model fine-tuning and tokenization.
txt, textLicense: CC-BY-4.0A collection of Burmese sentences curated with numerical classifiers, digits, and measurement units across various registers and linguistic styles.
txt, textLicense: CC-BY-4.0A handcrafted conversational dialogue dataset designed to fine-tune Burmese large language models into a specific romantic roleplay persona.
jsonl, textLicense: CC-BY-NC-4.0A parallel text corpus aligning Pali text with Burmese translations, romanized transliterations, and grammatical context notes.
csvLicense: CC-BY-NC-4.0An image-to-text dataset containing rendered images of ancient Burmese Pyu characters along with structural labels for script recognition research.
parquet, image, textLicense: CC-BY-NC-4.0A comprehensive collection of Burmese news articles crawled from Voice of America focusing on high-quality news content.
jsonl, textLicense: CC-BY-4.0A high-quality structured collection of the Buddhist Pali Canon transcribed in the Myanmar script covering canonical texts, commentaries, and sub-commentaries.
jsonl, textLicense: CC-BY-4.0A high-volume synthetically augmented dataset of Burmese word formations following strict morphological and grammatical rules for structural analysis.
jsonl, textLicense: CC-BY-4.0A human-curated dataset of code-mixed Burmese and English sentences reflecting contemporary linguistic patterns of Myanmar's digital society.
txt, textLicense: CC-BY-4.0An extensive and highly structured classical-to-modern language mapping dictionary dataset designed for instruction tuning and machine translation.
jsonl, textLicense: CC-BY-4.0A specialized binary text classification dataset for distinguishing between formal written style and informal spoken style Burmese text.
csv, textLicense: CC-BY-4.0A curated hybrid image dataset combining human-annotated handwritten samples and computer-generated font variations of native Burmese digits for OCR.
parquet, image, textLicense: CC-BY-4.0A massive-scale synthetic image dataset representing the exhaustive structural combinatorial matrix of syllables and diacritic stacks in the Burmese script.
parquet, image, textLicense: CC-BY-4.0A vocabulary-based synthetic image dataset mapping authentic and meaningful words extracted from Myanmar Wiktionary for real-world OCR tasks.
parquet, image, textLicense: CC-BY-4.0A parallel noisy-to-clean text dataset focusing on spelling errors, phonetic confusions, and keyboard typos in the Burmese language.
csv, textLicense: CC-BY-4.0A human-curated Burmese question-answering dataset focused on planetary science and astronomy utilizing natural conversational polite speech.
jsonl, textLicense: CC-BY-4.0A mathematically exhaustive collection of synthetic Burmese pseudo-syllables structured systematically for stress-testing tokenizers and spell-checkers.
csv, textLicense: CC-BY-4.0A cleaned, processed, and high-quality collection of Burmese Wikipedia articles formatted with sentence segmentation and syllable tokenization columns.
jsonl, textLicense: CC-BY-4.0A refined monolingual collection strictly filtered to ensure sentences contain exclusively Burmese script blocks by eliminating cross-lingual noise.
jsonl, textLicense: CC-BY-4.0A conversational question-and-answering dataset compiling geographical, historical, and cultural data about various cities in Myanmar using spoken tone.
jsonl, textLicense: CC-BY-4.0A granular sentence-by-sentence text dataset systematically parsed from the official articles of the Ministry of Information website for language modeling.
parquet, textLicense: CC-BY-4.0A collection of full-length official news and announcement articles extracted directly from the Ministry of Information portal for Myanmar NLP growth.
parquet, textLicense: CC-BY-4.0A high-quality parallel corpus mapping formal written literary Burmese directly to daily informal spoken colloquial Burmese for style transfer.
csv, textLicense: CC-BY-4.0An exhaustive mapping dataset bridging Arabic numerals and native Burmese digit text strings across short and full full-text contextual styles.
csv, textLicense: CC-BY-4.0A single-label text classification dataset framing high-register Burmese prose and literary emotions through the nine traditional aesthetics.
csv, textLicense: CC-BY-4.0A high-fidelity audio dataset pairing synthesized male and female voices with verified texts to advance Text-to-Speech systems for Burmese.
parquet, audio, textLicense: CC-BY-4.0A parallel dataset mapping native Burmese script to its corresponding Latin/Myanglish phonetic transliterations widely used in casual online chat.
csv, textLicense: CC-BY-4.0A large-scale open public-domain Burmese speech dataset curated from public-service educational broadcasts.
tar, mp3, jsonLicense: CC0-1.0A collection of Burmese news article snippets categorized into Business, Entertainment, Politics, and Sport.
csv, textLicense: GPL-3.0A human-translated and machine-translated Burmese corpus extending the global XNLI framework for text inference tasks.
csv, textLicense: CC-BY-NC-4.0A crowdsourced high-quality Burmese speech dataset extracted and filtered from the multilingual OpenSLR index.
wav, audio, textLicense: CC-BY-SA-4.0A massive collection of high-quality written Burmese sentences derived from web data to support versatile language modeling.
parquet, textLicense: CC0-1.0A large-scale corpus of filtered spoken Burmese sentences optimized for acoustic text-to-speech models and speech processing.
parquet, textLicense: CC-BY-4.0A synthetic image-to-text dataset containing rendered Burmese text with various styles, blurs, and distortions for training OCR models.
image, text, paddleocrLicense: MITA preprocessed subset of the CC100 corpus containing pure Burmese text converted uniformly from Zawgyi to standard Unicode.
text, datasetLicense: CULTURAXA detailed Burmese-to-Burmese word dictionary providing linguistic properties like alphabet, phonetics, parts of speech, and etymology roots.
parquet, textLicense: MITA collection of anonymized and normalized consumer product reviews sourced from e-commerce communities using Burmese Unicode.
csv, textLicense: MITA structured syntactic dataset demonstrating rule-based sentence scaling across word lengths generated from a single grammar structure.
csv, textLicense: MITA richly structured and fully aligned textual dataset of the Myanmar Bible New World Translation for text tasks.
csv, textLicense: CUSTOMA parallel corpus of high-quality audio recordings and matching Burmese transcriptions covering selected books of the Bible.
soundfolder, audio, textLicense: CUSTOMAn expanded parallel corpus linking high-quality Bible narration audio files with faithful Burmese text transcriptions.
soundfolder, audio, textLicense: CUSTOMAn audio-transcript parallel corpus extracted from public YouTube videos published by the official PVTV media channel.
webdataset, audio, textLicense: CC0-1.0Sentence-level audio chunks and metadata extracted from Voice of America Burmese morning radio programs for speech modeling.
webdataset, audio, textLicense: PDDLA massive historical unsupervised audio corpus of sentence-level chunks spanning a decade of VOA Burmese radio broadcasts.
webdataset, audio, textLicense: PDDLA comprehensive bilingual English and Myanmar gazetteer mapping unique placenames for villages, townships, and regions.
csv, textLicense: CC-BY-NC-SA-4.0An English-Myanmar parallel translation corpus focused on natural spoken-style phrasing curated for machine translation research.
csv, textLicense: CC-BY-NC-SA-4.0A token classification dataset for Burmese syllable and chunk segmentation formatted using the sequence labeling BIO scheme.
parquetLicense: CC-BY-4.0A preprocessed and cleaned sequence labeling token classification dataset annotated with standard Myanmar part-of-speech tags.
parquetLicense: CC-BY-4.0A linguistically enriched collection of traditional Burmese idioms providing dual-register explanations and contextual narrative stories.
jsonl, textLicense: CC0-1.0A collection of Arabic names paired with corresponding English and Myanmar script transliterations for entity recognition tasks.
json, textLicense: CC0-1.0A high-quality multi-parallel translation corpus aligning Arabic text with trusted human scholarship and multiple AI realizations.
jsonl, textLicense: CC-BY-NC-4.0A digitized multi-language dictionary based on the lexicographical work of U Hote Sein for Buddhist and Pali studies.
json, textLicense: CUSTOMA master index of unique Pali words rendered and normalized in standard Myanmar Unicode script for computational workflows.
json, textLicense: CC0-1.0A multi-lingual parallel text translation dataset linking Myanmar, English, and Korean languages.
parquet, textLicense: MITA multi-lingual parallel text translation dataset linking Myanmar, English, and Chinese languages.
parquet, textLicense: MITA multi-lingual parallel text translation dataset linking Myanmar, English, and Japanese languages.
parquet, textLicense: MITA smaller focused multilingual text translation dataset containing parallel entries in English, Myanmar, and Chinese.
csv, textLicense: MITA cleaned, Unicode-normalized, and length-sorted version of the Burmese translation subset from FineTranslations.
parquet, textLicense: CC0-1.0The complete Myanmar translation of the Pali Canon, commentaries, and historical treatises converted into a clean structured schema.
jsonl, textLicense: CC0-1.0An informational public health question-answering dataset about the Monkeypox virus compiled in the Burmese language.
csv, textLicense: CC-BY-4.0Crop production, market prices, agri-household surveys from IFPRI Myanmar program.
CSV, STATA, XLSXLicense: Open (with registration)Official crop calendar with planting and harvest seasons for major crops.
PDF, CSVLicense: OpenAn instruction-tuning dataset containing question-answer pairs about crops, techniques, and chemicals in the Burmese language.
csv, textLicense: CC-BY-4.0Myanmar news articles in Business, Entertainment, Politics, Sport categories.
CSVLicense: GPL-3.0Collection of Myanmar news articles from BBC, VOA, DVB etc.
ParquetLicense: MITMyanmar-English parallel news translation dataset.
UnknownLicense: UnknownMyanmar-specific hate speech datasets and experiments.
VariousLicense: UnknownMyanmar social media sentiment analysis dataset.
UnknownLicense: UnknownBurmese news category dataset from DVB.
CSVLicense: UnknownUnknownLicense: UnknownNo datasets matching your criteria.