A robust, regression‑tested pipeline for retrieving publication counts and recent publications for educational institutions using the OpenAlex API.
This project supports two modes:
- Level 1 — Summary CSV with publication counts
- Level 2 — Detailed Markdown with recent publications (sample of N)
The system includes:
- Automatic retry
- Polite rate limiting
- Parallel requests
- Caching
- Fallback institution name resolution
- Regression test suite
- Golden file refresh tool
OpenAlex/
│
├── get_counts.py
├── regression_test.py
├── refresh_golden.py
│
├── input_csv.csv
├── regression_test.csv
│
├── regression_test_expected_level1.csv
├── regression_test_expected_level2_structure.md
│
├── output_dir/ (normal output)
└── regression_output/ (regression test output)
| File | Purpose |
|---|---|
| get_counts.py | Main pipeline: Level 1 + Level 2 processing |
| regression_test.py | Runs regression tests for Level 1 and Level 2 |
| refresh_golden.py | Regenerates golden files automatically |
| regression_test.csv | Small 3‑row test dataset (MIT, Babson, Wellesley) |
| regression_test_expected_level1.csv | Golden file for Level 1 |
| regression_test_expected_level2_structure.md | Structure‑only golden file for Level 2 |
| output_dir/ | Normal output directory |
| regression_output/ | Output generated during regression tests |
python get_counts.py --level 1 --input input_csv.csv --output_dir output_dir
Produces:
output_dir/publication_counts.csv
python get_counts.py --level 2 --input input_csv.csv --output_dir output_dir
Produces:
output_dir/detailed_publications.md
Level 2 retrieves only the most recent N publications (default: 20).
The regression suite ensures that:
- Level 1 publication counts match the golden file
- Level 2 output has the correct structure
python regression_test.py
You will see:
Level 1: PASS
Level 2: PASS
OpenAlex is a live database — publication counts change daily.
When counts drift, the regression test will fail (correctly).
To update the golden files:
python refresh_golden.py
This regenerates:
regression_test_expected_level1.csvregression_test_expected_level2_structure.md
Then re‑run:
python regression_test.py
If OpenAlex cannot find an institution using the org_name field:
- The pipeline extracts the domain from
WebsiteAddress - It infers a simpler name (e.g.,
home.dartmouth.edu→ “Dartmouth”) - It retries the OpenAlex search
This fixes cases like:
- “TRUSTEES OF DARTMOUTH COLLEGE”
- “PRESIDENT AND FELLOWS OF HARVARD COLLEGE”
- “THE CORPORATION OF BROWN UNIVERSITY”
At the top of get_counts.py:
LEVEL = 1
RECENT_PUBLICATIONS = 20
MAX_WORKERS = 4You can change:
- default level
- number of recent publications
- parallelism
Command‑line arguments override the default level.
- Edit your input CSV
- Run Level 1 or Level 2
- When OpenAlex changes and regression tests fail:
python refresh_golden.py python regression_test.py - Commit both golden files and the updated code
This keeps your pipeline stable and reproducible.