Build unified Sinhala / Tamil / English corpora for fine-tuning Phi-4 14B (Microsoft). Pipeline: scrapers → extract MADLAD → normalize → merge & deduplicate.
- Repo is ready to push. Large data and outputs are ignored via
.gitignore(data/,corpora/,output/,*.parquet, scraper caches). - Only code and small config are committed; datasets live in Kaggle Datasets or local disk.
- New Notebook on Kaggle.
- Add datasets (e.g. MADLAD_CulturaX_cleaned or your copy; optional: add a dataset with
raw/sinhala_3M.txt,raw/tamil_3M.txt). - Clone this repo or upload the
dataset_pipelinefolder and set the repo as the working directory. - Set env (in the notebook or “Settings” → “Environment”):
DATA_DIR=/kaggle/input/<your-dataset>/(or/kaggle/working/data– all datasets live underdata/)OUTPUT_DIR=/kaggle/working/dataCLEANUP_WORK=1to delete intermediatework/after merge (saves disk)- Optional:
SKIP_MADLAD=1if you only have raw.txtand no parquet. - Optional:
LIMIT_SI=500000,LIMIT_TA=500000,LIMIT_EN=500000to cap lines per language (saves disk).
- Install (one cell):
!pip install -r dataset_pipeline/requirements.txt - Run pipeline (one cell):
!python dataset_pipeline/run_pipeline.py
Outputs:unified_si.txt,unified_ta.txt,unified_en.txtunderOUTPUT_DIR.
Do not run the NIE or Sri Lanka scrapers on Kaggle; run them locally and, if needed, add their outputs to DATA_DIR/raw/ or your Kaggle dataset.
Runs all scrapers (NIE PDFs, Sri Lanka 100+ sources), your 3M lines, MADLAD, synthetic, UD Sinhala, then the full corpus pipeline. Target model: Phi-4 14B.
1. Add data – Attach your Kaggle dataset(s) with:
raw/sinhala_3M.txt,raw/tamil_3M.txt(required)raw/sinhala_sentences.txt,raw/tamil_sentences.txt,raw/ud_sinhala_sentences.txt(optional)MADLAD_CulturaX_cleaned/data/*.parquet(optional)
2. New Notebook – Paste and run:
# Cell 1: Clone repo
!git clone https://github.com/Januth1234/Orin-X.git
%cd Orin-X
# Cell 2: Run full pipeline (scrapers + corpus)
!python kaggle_run_all.pyOutputs: unified_si.txt, unified_ta.txt, unified_en.txt in data/ (or /kaggle/working/data/ on Kaggle). Use these for fine-tuning Phi-4 14B (QLoRA) in a follow-up notebook.
Using your own datasets with scraped data: Put your unified_*.txt (or raw .txt) in the same dataset or DATA_DIR. The pipeline merges everything; add your files to raw/ (e.g. raw/my_sinhala.txt) or use a Kaggle dataset that already contains unified_si.txt / unified_ta.txt. For fine-tuning, point run_finetune.py at the folder that has the merged unified_*.txt (scraped + yours).
Note: Scrapers use polite delays (~2.5s between requests). NIE + Sri Lanka can take 1–3 hours. Enable "Internet" in Notebook settings.
After you have unified_*.txt (from scrapers + your own data), run QLoRA in one go:
python run_finetune.py
# Or on Kaggle (data path = your dataset with unified_*.txt):
python run_finetune.py --data-path /kaggle/input/.../your-dataset --output-dir /kaggle/working/phi4-qloraUses 4-bit quant, LoRA, balanced 2× GPU + CPU offload. See run_finetune.py for --max-samples, --epochs, etc.
On a TPU VM (e.g. v5e-8, 330 GB RAM), use FSDP + bf16 (no 4-bit; BitsAndBytes is GPU-only):
export PJRT_DEVICE=TPU # if not already set
python run_finetune_tpu.py --data-path /path/to/unified_txt --output-dir /path/out --max-samples 500000 --epochs 1Requires: optimum-tpu (install from Google’s libtpu index), PyTorch/XLA, same unified_*.txt data. See run_finetune_tpu.py for options.
| Step | Script | Purpose |
|---|---|---|
| 1 | extract_madlad_to_txt.py |
Stream MADLAD parquet → per-language .txt (si/ta/en) |
| 2 | normalize_txt.py |
NFC + whitespace norm, min/max length |
| 3 | merge_dedup.py |
Merge per-language files, dedupe by line hash → unified_<lang>.txt |
Single entry point: dataset_pipeline/run_pipeline.py (uses DATA_DIR, OUTPUT_DIR, WORK_DIR).
dataset_pipeline/
run_pipeline.py # entry point
extract_madlad_to_txt.py
normalize_txt.py
merge_dedup.py
nie_scraper/ # run locally only
sri_lanka_scraper/ # run locally only
Data stays out of the repo (.gitignore). On Kaggle, point DATA_DIR to input datasets and OUTPUT_DIR to /kaggle/working.
CLEANUP_WORK=1– deleteswork/(extracted + normalized) after merge; keeps onlyunified_*.txt.- Attach datasets – use “Add data” so MADLAD and raw
.txtlive in/kaggle/input/(read-only); no copy to working disk. LIMIT_SI/LIMIT_TA/LIMIT_EN– cap unique lines per language (e.g.500000) to shrink output.SKIP_MADLAD=1– if you only have raw.txtin your dataset, skip MADLAD extraction entirely.