Skip to content

Repository files navigation

Orin X – Trilingual corpus pipeline

Build unified Sinhala / Tamil / English corpora for fine-tuning Phi-4 14B (Microsoft). Pipeline: scrapers → extract MADLAD → normalize → merge & deduplicate.

Push to GitHub

  • Repo is ready to push. Large data and outputs are ignored via .gitignore (data/, corpora/, output/, *.parquet, scraper caches).
  • Only code and small config are committed; datasets live in Kaggle Datasets or local disk.

Run on Kaggle (without filling storage)

  1. New Notebook on Kaggle.
  2. Add datasets (e.g. MADLAD_CulturaX_cleaned or your copy; optional: add a dataset with raw/sinhala_3M.txt, raw/tamil_3M.txt).
  3. Clone this repo or upload the dataset_pipeline folder and set the repo as the working directory.
  4. Set env (in the notebook or “Settings” → “Environment”):
    • DATA_DIR=/kaggle/input/<your-dataset>/ (or /kaggle/working/data – all datasets live under data/)
    • OUTPUT_DIR=/kaggle/working/data
    • CLEANUP_WORK=1 to delete intermediate work/ after merge (saves disk)
    • Optional: SKIP_MADLAD=1 if you only have raw .txt and no parquet.
    • Optional: LIMIT_SI=500000, LIMIT_TA=500000, LIMIT_EN=500000 to cap lines per language (saves disk).
  5. Install (one cell):
    !pip install -r dataset_pipeline/requirements.txt
  6. Run pipeline (one cell):
    !python dataset_pipeline/run_pipeline.py
    Outputs: unified_si.txt, unified_ta.txt, unified_en.txt under OUTPUT_DIR.

Do not run the NIE or Sri Lanka scrapers on Kaggle; run them locally and, if needed, add their outputs to DATA_DIR/raw/ or your Kaggle dataset.

Full Kaggle run (scrapers + pipeline → Phi-4 14B)

Runs all scrapers (NIE PDFs, Sri Lanka 100+ sources), your 3M lines, MADLAD, synthetic, UD Sinhala, then the full corpus pipeline. Target model: Phi-4 14B.

1. Add data – Attach your Kaggle dataset(s) with:

  • raw/sinhala_3M.txt, raw/tamil_3M.txt (required)
  • raw/sinhala_sentences.txt, raw/tamil_sentences.txt, raw/ud_sinhala_sentences.txt (optional)
  • MADLAD_CulturaX_cleaned/data/*.parquet (optional)

2. New Notebook – Paste and run:

# Cell 1: Clone repo
!git clone https://github.com/Januth1234/Orin-X.git
%cd Orin-X

# Cell 2: Run full pipeline (scrapers + corpus)
!python kaggle_run_all.py

Outputs: unified_si.txt, unified_ta.txt, unified_en.txt in data/ (or /kaggle/working/data/ on Kaggle). Use these for fine-tuning Phi-4 14B (QLoRA) in a follow-up notebook.

Using your own datasets with scraped data: Put your unified_*.txt (or raw .txt) in the same dataset or DATA_DIR. The pipeline merges everything; add your files to raw/ (e.g. raw/my_sinhala.txt) or use a Kaggle dataset that already contains unified_si.txt / unified_ta.txt. For fine-tuning, point run_finetune.py at the folder that has the merged unified_*.txt (scraped + yours).

Note: Scrapers use polite delays (~2.5s between requests). NIE + Sri Lanka can take 1–3 hours. Enable "Internet" in Notebook settings.

One-command fine-tune (Phi-4-reasoning-vision-15B, 2× T4, 30 GB RAM)

After you have unified_*.txt (from scrapers + your own data), run QLoRA in one go:

python run_finetune.py
# Or on Kaggle (data path = your dataset with unified_*.txt):
python run_finetune.py --data-path /kaggle/input/.../your-dataset --output-dir /kaggle/working/phi4-qlora

Uses 4-bit quant, LoRA, balanced 2× GPU + CPU offload. See run_finetune.py for --max-samples, --epochs, etc.

Fine-tune on TPU (v5e-8, Google Cloud)

On a TPU VM (e.g. v5e-8, 330 GB RAM), use FSDP + bf16 (no 4-bit; BitsAndBytes is GPU-only):

export PJRT_DEVICE=TPU   # if not already set
python run_finetune_tpu.py --data-path /path/to/unified_txt --output-dir /path/out --max-samples 500000 --epochs 1

Requires: optimum-tpu (install from Google’s libtpu index), PyTorch/XLA, same unified_*.txt data. See run_finetune_tpu.py for options.

Pipeline steps (what runs)

Step Script Purpose
1 extract_madlad_to_txt.py Stream MADLAD parquet → per-language .txt (si/ta/en)
2 normalize_txt.py NFC + whitespace norm, min/max length
3 merge_dedup.py Merge per-language files, dedupe by line hash → unified_<lang>.txt

Single entry point: dataset_pipeline/run_pipeline.py (uses DATA_DIR, OUTPUT_DIR, WORK_DIR).

Layout (minimal for Kaggle)

dataset_pipeline/
  run_pipeline.py      # entry point
  extract_madlad_to_txt.py
  normalize_txt.py
  merge_dedup.py
nie_scraper/           # run locally only
sri_lanka_scraper/     # run locally only

Data stays out of the repo (.gitignore). On Kaggle, point DATA_DIR to input datasets and OUTPUT_DIR to /kaggle/working.

Max storage savings on Kaggle

  • CLEANUP_WORK=1 – deletes work/ (extracted + normalized) after merge; keeps only unified_*.txt.
  • Attach datasets – use “Add data” so MADLAD and raw .txt live in /kaggle/input/ (read-only); no copy to working disk.
  • LIMIT_SI / LIMIT_TA / LIMIT_EN – cap unique lines per language (e.g. 500000) to shrink output.
  • SKIP_MADLAD=1 – if you only have raw .txt in your dataset, skip MADLAD extraction entirely.

About

No description, website, or topics provided.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Used by

Contributors

Languages