Search Data Lam Hnyin Articles

Open Resource Registry

Myanmar Open Datasets Directory

A centralized, open-access catalog of structured datasets covering Myanmar NLP, geospatial mapping, humanitarian metrics, and technical benchmarks. Maintained by the community for researchers and machine learning practitioners.

Showing 76 datasetsCatalog ID: src/dataset.csv
DLH-NLP-001NLP & Language

Myanmar Written-Spoken Text Pairs

A parallel text dataset pairing formal written Myanmar sentences with their natural spoken alternatives for style transfer and NLP tasks.

#parallel-corpus#text#text-normalization#style-transfer#burmese
Format: csvLicense: CC-BY-NC-SA-4.0
By Khant Sint Heinn (Kalix Louis)
DLH-NLP-002NLP & Language

HFcourse-English-Burmese-Parallel-Corpus

An aligned English-Burmese parallel sentence corpus extracted from Hugging Face Course video subtitles to support machine translation research.

#parallel-corpus#text#machine-translation#subtitles#burmese
Format: csvLicense: CC-BY-4.0
By Khant Sint Heinn (Kalix Louis)
DLH-NLP-003NLP & Language

Myanmar-English General Text Translation

A high-quality, human-verified parallel corpus covering general domain texts, everyday conversations, and literature for Myanmar-English machine translation.

#parallel-corpus#text#machine-translation#general-domain#burmese
Format: parquetLicense: CC-BY-NC-SA-4.0
By Khant Sint Heinn (Kalix Louis)
DLH-NLP-004NLP & Language

Burmese Text Corpus For Natural Language Processing

A curated text corpus of grammatically sound Burmese sentences collected from news and educational websites for general NLP tasks.

#text-corpus#text#web-scraping#nlp#burmese
Format: jsonl, textLicense: CC-BY-4.0
By Khant Sint Heinn (Kalix Louis)
DLH-NLP-005NLP & Language

MyanmarOCR-ImageText Dataset

A diverse synthetic image-to-text dataset featuring Burmese characters rendered in various visual styles, colors, and backgrounds for OCR training.

#ocr#image-to-text#multimodal#synthetic-data#burmese
Format: parquet, image, textLicense: CC-BY-4.0
By Khant Sint Heinn (Kalix Louis)
DLH-NLP-006NLP & Language

Myanmar Linguistic Ambiguities - Dataset 001

A labeled corpus designed for grammatical error correction and disambiguation of the frequently confused Burmese particles 'လဲ' and 'လည်း'.

#grammar-correction#text#classification#linguistics#burmese
Format: csvLicense: CC-BY-4.0
By Khant Sint Heinn (Kalix Louis)
DLH-NLP-007NLP & Language

General Burmese Sentences

An open-source text corpus containing general-domain sentences written in natural Burmese to support language model fine-tuning and tokenization.

#text-corpus#text#language-modeling#fine-tuning#burmese
Format: txt, textLicense: CC-BY-4.0
By Khant Sint Heinn (Kalix Louis)
DLH-NLP-008NLP & Language

Myanmar General Numerals Corpus

A collection of Burmese sentences curated with numerical classifiers, digits, and measurement units across various registers and linguistic styles.

#text-corpus#text#numerals#classifiers#linguistics#burmese
Format: txt, textLicense: CC-BY-4.0
By Khant Sint Heinn (Kalix Louis)
DLH-NLP-009NLP & Language

Boyfriend-LLM-Instruct

A handcrafted conversational dialogue dataset designed to fine-tune Burmese large language models into a specific romantic roleplay persona.

#conversational-ai#text#instruct-dataset#roleplay#persona#burmese
Format: jsonl, textLicense: CC-BY-NC-4.0
By Khant Sint Heinn (Kalix Louis)
DLH-NLP-010NLP & Language

Pali-Myanmar Parallel Corpus 1K

A parallel text corpus aligning Pali text with Burmese translations, romanized transliterations, and grammatical context notes.

#parallel-corpus#text#pali#transliteration#linguistics
Format: csvLicense: CC-BY-NC-4.0
By Khant Sint Heinn (Kalix Louis)
DLH-NLP-011NLP & Language

Burmese Pyu Character Recognition

An image-to-text dataset containing rendered images of ancient Burmese Pyu characters along with structural labels for script recognition research.

#ancient-scripts#image-to-text#ocr#pyu#script-recognition
Format: parquet, image, textLicense: CC-BY-NC-4.0
By Khant Sint Heinn (Kalix Louis)
DLH-NLP-012NLP & Language

Burmese VOA News Dataset

A comprehensive collection of Burmese news articles crawled from Voice of America focusing on high-quality news content.

#news#text-corpus#nlp#science-technology#burmese
Format: jsonl, textLicense: CC-BY-4.0
By DatarrX org
DLH-NLP-013NLP & Language

Myanmar Tipitaka Pali Canon Dataset

A high-quality structured collection of the Buddhist Pali Canon transcribed in the Myanmar script covering canonical texts, commentaries, and sub-commentaries.

#pali#religion#text-corpus#linguistics#classical-language
Format: jsonl, textLicense: CC-BY-4.0
By DatarrX org
DLH-NLP-014NLP & Language

Burmese Morpho-Synthetic Word Formations Dataset

A high-volume synthetically augmented dataset of Burmese word formations following strict morphological and grammatical rules for structural analysis.

#synthetic-data#morphology#grammar#linguistics#burmese
Format: jsonl, textLicense: CC-BY-4.0
By DatarrX org
DLH-NLP-015NLP & Language

Burmese-English Code-Mixed Dataset

A human-curated dataset of code-mixed Burmese and English sentences reflecting contemporary linguistic patterns of Myanmar's digital society.

#code-mixing#slang#conversational-ai#social-media#burmese
Format: txt, textLicense: CC-BY-4.0
By DatarrX org
DLH-NLP-016NLP & Language

Pali-Myanmar Dictionary Dataset

An extensive and highly structured classical-to-modern language mapping dictionary dataset designed for instruction tuning and machine translation.

#dictionary#mapping#instruction-tuning#part-of-speech#pali
Format: jsonl, textLicense: CC-BY-4.0
By DatarrX org
DLH-NLP-017NLP & Language

Myanmar Text Style Classification Dataset

A specialized binary text classification dataset for distinguishing between formal written style and informal spoken style Burmese text.

#text-classification#style-detection#formal-informal#diglossia#burmese
Format: csv, textLicense: CC-BY-4.0
By DatarrX org
DLH-NLP-018NLP & Language

Myanmar Numeral Glyphs Dataset

A curated hybrid image dataset combining human-annotated handwritten samples and computer-generated font variations of native Burmese digits for OCR.

#ocr#digits#handwritten#image-classification#fonts
Format: parquet, image, textLicense: CC-BY-4.0
By DatarrX org
DLH-NLP-019NLP & Language

Myanmar Synthetic Syllable Glyphs Dataset

A massive-scale synthetic image dataset representing the exhaustive structural combinatorial matrix of syllables and diacritic stacks in the Burmese script.

#ocr#syllables#synthetic-data#diacritics#script-engineering
Format: parquet, image, textLicense: CC-BY-4.0
By DatarrX org
DLH-NLP-020NLP & Language

Myanmar Word Glyphs Vocabulary Dataset

A vocabulary-based synthetic image dataset mapping authentic and meaningful words extracted from Myanmar Wiktionary for real-world OCR tasks.

#ocr#vocabulary#wiktionary#synthetic-images#words
Format: parquet, image, textLicense: CC-BY-4.0
By DatarrX org
DLH-NLP-021NLP & Language

Myanmar Orthography Error Correction Dataset

A parallel noisy-to-clean text dataset focusing on spelling errors, phonetic confusions, and keyboard typos in the Burmese language.

#orthography#spelling-correction#typos#parallel-corpus#burmese
Format: csv, textLicense: CC-BY-4.0
By DatarrX org
DLH-NLP-022NLP & Language

Planet Facts Question Answering Dataset

A human-curated Burmese question-answering dataset focused on planetary science and astronomy utilizing natural conversational polite speech.

#question-answering#astronomy#science#conversational-ai#polite-tone
Format: jsonl, textLicense: CC-BY-4.0
By DatarrX org
DLH-NLP-023NLP & Language

Myanmar Synthetic Pseudo-Syllables Dataset

A mathematically exhaustive collection of synthetic Burmese pseudo-syllables structured systematically for stress-testing tokenizers and spell-checkers.

#synthetic-data#syllables#stress-testing#tokenization#edge-cases
Format: csv, textLicense: CC-BY-4.0
By DatarrX org
DLH-NLP-024NLP & Language

Myanmar Wikipedia Text Dataset

A cleaned, processed, and high-quality collection of Burmese Wikipedia articles formatted with sentence segmentation and syllable tokenization columns.

#wikipedia#text-corpus#sentence-segmentation#tokenization#burmese
Format: jsonl, textLicense: CC-BY-4.0
By DatarrX org
DLH-NLP-025NLP & Language

Myanmar Monolingual Pure Wikipedia Dataset

A refined monolingual collection strictly filtered to ensure sentences contain exclusively Burmese script blocks by eliminating cross-lingual noise.

#wikipedia#monolingual#pure-text#text-filtering#burmese
Format: jsonl, textLicense: CC-BY-4.0
By DatarrX org
DLH-NLP-026NLP & Language

Myanmar Cities Question Answering Dataset

A conversational question-and-answering dataset compiling geographical, historical, and cultural data about various cities in Myanmar using spoken tone.

#question-answering#geography#history#chatbot#conversational-ai
Format: jsonl, textLicense: CC-BY-4.0
By DatarrX org
DLH-NLP-027NLP & Language

Ministry of Information Article Sentences Dataset

A granular sentence-by-sentence text dataset systematically parsed from the official articles of the Ministry of Information website for language modeling.

#news#formal-text#sentences#language-modeling#official-articles
Format: parquet, textLicense: CC-BY-4.0
By DatarrX org
DLH-NLP-028NLP & Language

Ministry of Information Full Articles Dataset

A collection of full-length official news and announcement articles extracted directly from the Ministry of Information portal for Myanmar NLP growth.

#news#full-text#articles#official-announcements#text-corpus
Format: parquet, textLicense: CC-BY-4.0
By DatarrX org
DLH-NLP-029NLP & Language

Myanmar Written-Spoken Parallel Dataset

A high-quality parallel corpus mapping formal written literary Burmese directly to daily informal spoken colloquial Burmese for style transfer.

#parallel-corpus#formal-informal#diglossia#style-transfer#burmese
Format: csv, textLicense: CC-BY-4.0
By DatarrX org
DLH-NLP-030NLP & Language

Burmese Numbers Normalization Dataset

An exhaustive mapping dataset bridging Arabic numerals and native Burmese digit text strings across short and full full-text contextual styles.

#text-normalization#numerals#digits#counting#text-to-speech
Format: csv, textLicense: CC-BY-4.0
By DatarrX org
DLH-NLP-031NLP & Language

Nava-Rasa Literary Aesthetics Text Dataset

A single-label text classification dataset framing high-register Burmese prose and literary emotions through the nine traditional aesthetics.

#text-classification#literature#emotions#sentiment-analysis#aesthetics
Format: csv, textLicense: CC-BY-4.0
By DatarrX org
DLH-NLP-032NLP & Language

Burmese Synthetic Speech Audio Dataset

A high-fidelity audio dataset pairing synthesized male and female voices with verified texts to advance Text-to-Speech systems for Burmese.

#text-to-speech#audio#speech-synthesis#audio-dataset#voice-persona
Format: parquet, audio, textLicense: CC-BY-4.0
By DatarrX org
DLH-NLP-033NLP & Language

Myanmar to Myanglish Romanization Parallel Dataset

A parallel dataset mapping native Burmese script to its corresponding Latin/Myanglish phonetic transliterations widely used in casual online chat.

#transliteration#myanglish#parallel-corpus#chat-text#romanization
Format: csv, textLicense: CC-BY-4.0
By DatarrX org
DLH-NLP-034NLP & Language

NUG Myanmar ASR Dataset

A large-scale open public-domain Burmese speech dataset curated from public-service educational broadcasts.

#asr#speech#audio#text#education#public-domain
Format: tar, mp3, jsonLicense: CC0-1.0
By freococo
DLH-NLP-035NLP & Language

Myanmar News Classification Dataset

A collection of Burmese news article snippets categorized into Business, Entertainment, Politics, and Sport.

#news#classification#text#articles#dataset
Format: csv, textLicense: GPL-3.0
By Aye Hninn Khine
DLH-NLP-036NLP & Language

Myanmar Cross-lingual Natural Language Inference Dataset

A human-translated and machine-translated Burmese corpus extending the global XNLI framework for text inference tasks.

#nli#inference#text#translation#parallel-corpus
Format: csv, textLicense: CC-BY-NC-4.0
By Aung Kyaw Htet
DLH-NLP-037NLP & Language

Myanmar Speech Crowdsourced Audio Dataset

A crowdsourced high-quality Burmese speech dataset extracted and filtered from the multilingual OpenSLR index.

#asr#speech#audio#crowdsourced#text-transcripts
Format: wav, audio, textLicense: CC-BY-SA-4.0
By Chuu Htet Naing
DLH-NLP-038NLP & Language

Myanmar Written Text Corpus

A massive collection of high-quality written Burmese sentences derived from web data to support versatile language modeling.

#text-corpus#text#large-scale#nlp#web-data
Format: parquet, textLicense: CC0-1.0
By freococo
DLH-NLP-039NLP & Language

Myanmar Spoken Language Corpus

A large-scale corpus of filtered spoken Burmese sentences optimized for acoustic text-to-speech models and speech processing.

#text-corpus#spoken-text#speech-processing#text#open-data
Format: parquet, textLicense: CC-BY-4.0
By freococo
DLH-NLP-040NLP & Language

Myanmar Synthetic OCR Image Dataset

A synthetic image-to-text dataset containing rendered Burmese text with various styles, blurs, and distortions for training OCR models.

#ocr#image-to-text#synthetic-data#text-recognition#fonts
Format: image, text, paddleocrLicense: MIT
By Chuu Htet Naing
DLH-NLP-041NLP & Language

Myanmar Standardized CC100 Clean Dataset

A preprocessed subset of the CC100 corpus containing pure Burmese text converted uniformly from Zawgyi to standard Unicode.

#text-corpus#text#unicode-conversion#web-data#preprocessing
Format: text, datasetLicense: CULTURAX
By Chuu Htet Naing
DLH-NLP-042NLP & Language

Burmese Lexical Dictionary Dataset

A detailed Burmese-to-Burmese word dictionary providing linguistic properties like alphabet, phonetics, parts of speech, and etymology roots.

#dictionary#linguistics#vocabulary#part-of-speech#words
Format: parquet, textLicense: MIT
By Pyae Sone Myo
DLH-NLP-043NLP & Language

Burmese E-Commerce Consumer Reviews Dataset

A collection of anonymized and normalized consumer product reviews sourced from e-commerce communities using Burmese Unicode.

#reviews#e-commerce#sentiment-analysis#customer-feedback#text
Format: csv, textLicense: MIT
By Pyae Sone Myo
DLH-NLP-044NLP & Language

Burmese Single Grammar Pattern Sentence Dataset

A structured syntactic dataset demonstrating rule-based sentence scaling across word lengths generated from a single grammar structure.

#grammar#linguistics#rule-based#syntax#sentence-generation
Format: csv, textLicense: MIT
By freococo
DLH-NLP-045NLP & Language

Jehovah's Witnesses Myanmar Bible Dataset

A richly structured and fully aligned textual dataset of the Myanmar Bible New World Translation for text tasks.

#bible#religion#text-generation#text-classification#summarization
Format: csv, textLicense: CUSTOM
By freococo
DLH-NLP-046NLP & Language

Jehovah's Witnesses Myanmar Bible 10-Books Audio Dataset

A parallel corpus of high-quality audio recordings and matching Burmese transcriptions covering selected books of the Bible.

#asr#speech#audio#bible#transcription
Format: soundfolder, audio, textLicense: CUSTOM
By freococo
DLH-NLP-047NLP & Language

Jehovah's Witnesses Myanmar Bible 56-Books Audio Dataset

An expanded parallel corpus linking high-quality Bible narration audio files with faithful Burmese text transcriptions.

#asr#speech#audio#bible#text-to-speech
Format: soundfolder, audio, textLicense: CUSTOM
By freococo
DLH-NLP-048NLP & Language

PVTV Myanmar Audio Speech Recognition Dataset

An audio-transcript parallel corpus extracted from public YouTube videos published by the official PVTV media channel.

#asr#audio-classification#speech#media#national-unity-government
Format: webdataset, audio, textLicense: CC0-1.0
By freococo
DLH-NLP-049NLP & Language

VOA Burmese Morning Radio Segmented Audio Dataset 2

Sentence-level audio chunks and metadata extracted from Voice of America Burmese morning radio programs for speech modeling.

#asr#audio-to-audio#audio-classification#radio#voice-of-america
Format: webdataset, audio, textLicense: PDDL
By freococo
DLH-NLP-050NLP & Language

VOA Burmese Morning Radio Segmented Audio Dataset 1

A massive historical unsupervised audio corpus of sentence-level chunks spanning a decade of VOA Burmese radio broadcasts.

#asr#radio#speech-modeling#audio-to-audio#audio-classification
Format: webdataset, audio, textLicense: PDDL
By freococo
DLH-NLP-051NLP & Language

Myanmar Bilingual Placename Dictionary Dataset

A comprehensive bilingual English and Myanmar gazetteer mapping unique placenames for villages, townships, and regions.

#placenames#gazetteer#dictionary#named-entity-recognition#bilingual
Format: csv, textLicense: CC-BY-NC-SA-4.0
By freococo
DLH-NLP-052NLP & Language

English-Myanmar Spoken Style Parallel Translation Dataset

An English-Myanmar parallel translation corpus focused on natural spoken-style phrasing curated for machine translation research.

#parallel-corpus#machine-translation#spoken-style#generative-ai#bilingual
Format: csv, textLicense: CC-BY-NC-SA-4.0
By freococo
DLH-NLP-053NLP & Language

Myanmar Text Chunk Segmentation Dataset

A token classification dataset for Burmese syllable and chunk segmentation formatted using the sequence labeling BIO scheme.

#token-classification#text-segmentation#bio-tagging#sequence-labeling#wikipedia
Format: parquetLicense: CC-BY-4.0
By Chuu Htet Naing
DLH-NLP-054NLP & Language

Myanmar Part-of-Speech Tagging Cleaned Dataset

A preprocessed and cleaned sequence labeling token classification dataset annotated with standard Myanmar part-of-speech tags.

#token-classification#part-of-speech#sequence-labeling#linguistic-annotations#mypos
Format: parquetLicense: CC-BY-4.0
By Chuu Htet Naing
DLH-NLP-055NLP & Language

Myanmar Idioms Lexicon Dataset

A linguistically enriched collection of traditional Burmese idioms providing dual-register explanations and contextual narrative stories.

#idioms#proverbs#culture#linguistics#bilingual
Format: jsonl, textLicense: CC0-1.0
By freococo
DLH-NLP-056NLP & Language

Arabic Names Transliteration Dataset

A collection of Arabic names paired with corresponding English and Myanmar script transliterations for entity recognition tasks.

#transliteration#names#named-entity-recognition#arabic#bilingual
Format: json, textLicense: CC0-1.0
By freococo
DLH-NLP-057NLP & Language

Myanmar Quran Parallel Translation Dataset

A high-quality multi-parallel translation corpus aligning Arabic text with trusted human scholarship and multiple AI realizations.

#parallel-corpus#quran#translation#human-vs-ai#evaluation
Format: jsonl, textLicense: CC-BY-NC-4.0
By freococo
DLH-NLP-058NLP & Language

Myanmar English Pali Dictionary Dataset

A digitized multi-language dictionary based on the lexicographical work of U Hote Sein for Buddhist and Pali studies.

#dictionary#pali#lexicography#buddhist-studies#multilingual
Format: json, textLicense: CUSTOM
By freococo
DLH-NLP-059NLP & Language

Pali Words in Myanmar Script Master Index

A master index of unique Pali words rendered and normalized in standard Myanmar Unicode script for computational workflows.

#vocabulary#master-index#pali#script-normalization#linguistics
Format: json, textLicense: CC0-1.0
By freococo
DLH-NLP-060NLP & Language

Myanmar English Korean Parallel Dataset

A multi-lingual parallel text translation dataset linking Myanmar, English, and Korean languages.

#parallel-corpus#machine-translation#multilingual#text
Format: parquet, textLicense: MIT
By Hmue Gyi
DLH-NLP-061NLP & Language

Myanmar English Chinese Parallel Dataset

A multi-lingual parallel text translation dataset linking Myanmar, English, and Chinese languages.

#parallel-corpus#machine-translation#multilingual#text
Format: parquet, textLicense: MIT
By Hmue Gyi
DLH-NLP-062NLP & Language

Myanmar English Japanese Parallel Dataset

A multi-lingual parallel text translation dataset linking Myanmar, English, and Japanese languages.

#parallel-corpus#machine-translation#multilingual#text
Format: parquet, textLicense: MIT
By Hmue Gyi
DLH-NLP-063NLP & Language

English Myanmar Chinese Multilingual Subset Dataset

A smaller focused multilingual text translation dataset containing parallel entries in English, Myanmar, and Chinese.

#parallel-corpus#multilingual#machine-translation#text
Format: csv, textLicense: MIT
By Hmue Gyi
DLH-NLP-064NLP & Language

Cleaned FineTranslations Myanmar English Dataset

A cleaned, Unicode-normalized, and length-sorted version of the Burmese translation subset from FineTranslations.

#translation#fine-tuning#web-data#unicode-normalization#curriculum-learning
Format: parquet, textLicense: CC0-1.0
By freococo
DLH-NLP-065NLP & Language

Myanmar Tipitaka Translation Canonical Books Dataset

The complete Myanmar translation of the Pali Canon, commentaries, and historical treatises converted into a clean structured schema.

#pali-canon#tipitaka#text-corpus#religion#dhamma
Format: jsonl, textLicense: CC0-1.0
By freococo
DLH-HLT-001Health

Mpox Monkeypox Informational Dataset

An informational public health question-answering dataset about the Monkeypox virus compiled in the Burmese language.

#health#public-education#monkeypox#text#question-answering
Format: csv, textLicense: CC-BY-4.0
By Min Si Thu
DLH-AGR-001Agriculture

Myanmar Agriculture Data (IFPRI MAPSA)

Crop production, market prices, agri-household surveys from IFPRI Myanmar program.

#Agriculture#Crop#Market Prices#Food
Format: CSV, STATA, XLSXLicense: Open (with registration)
By IFPRI / USAID
DLH-AGR-002Agriculture

Myanmar Crop Calendar (FAO GIEWS)

Official crop calendar with planting and harvest seasons for major crops.

#Crop Calendar#Season#Harvest#Planting
Format: PDF, CSVLicense: Open
By FAO GIEWS
DLH-AGR-003Agriculture

Myanmar Agriculture Question Answering Dataset

An instruction-tuning dataset containing question-answer pairs about crops, techniques, and chemicals in the Burmese language.

#agriculture#question-answering#instruction-tuning#text#rag
Format: csv, textLicense: CC-BY-4.0
By Min Si Thu
DLH-MED-001Media

ayehninnkhine/myanmar_news

Myanmar news articles in Business, Entertainment, Politics, Sport categories.

#News#Classification
Format: CSVLicense: GPL-3.0
By HuggingFace
DLH-MED-002Media

ThuraAung1601/myanmar-news

Collection of Myanmar news articles from BBC, VOA, DVB etc.

#News
Format: ParquetLicense: MIT
By HuggingFace
DLH-MED-004Media

ye-kyaw-thu/myHateSpeech

Myanmar-specific hate speech datasets and experiments.

#Hate Speech
Format: VariousLicense: Unknown
By GitHub