Skip to content

Repository files navigation

📘 README.md


OpenAlex Component of Pipeline

A robust, regression‑tested pipeline for retrieving publication counts and recent publications for educational institutions using the OpenAlex API.

This project supports two modes:

  • Level 1 — Summary CSV with publication counts
  • Level 2 — Detailed Markdown with recent publications (sample of N)

The system includes:

  • Automatic retry
  • Polite rate limiting
  • Parallel requests
  • Caching
  • Fallback institution name resolution
  • Regression test suite
  • Golden file refresh tool

📂 Project Structure

OpenAlex/
│
├── get_counts.py
├── regression_test.py
├── refresh_golden.py
│
├── input_csv.csv
├── regression_test.csv
│
├── regression_test_expected_level1.csv
├── regression_test_expected_level2_structure.md
│
├── output_dir/                 (normal output)
└── regression_output/          (regression test output)

Key Files

File Purpose
get_counts.py Main pipeline: Level 1 + Level 2 processing
regression_test.py Runs regression tests for Level 1 and Level 2
refresh_golden.py Regenerates golden files automatically
regression_test.csv Small 3‑row test dataset (MIT, Babson, Wellesley)
regression_test_expected_level1.csv Golden file for Level 1
regression_test_expected_level2_structure.md Structure‑only golden file for Level 2
output_dir/ Normal output directory
regression_output/ Output generated during regression tests

🚀 Running the Pipeline

Level 1 (publication counts)

python get_counts.py --level 1 --input input_csv.csv --output_dir output_dir

Produces:

output_dir/publication_counts.csv

Level 2 (recent publications)

python get_counts.py --level 2 --input input_csv.csv --output_dir output_dir

Produces:

output_dir/detailed_publications.md

Level 2 retrieves only the most recent N publications (default: 20).


🧪 Regression Testing

The regression suite ensures that:

  • Level 1 publication counts match the golden file
  • Level 2 output has the correct structure

Run the regression tests

python regression_test.py

You will see:

Level 1: PASS
Level 2: PASS

🔄 Refreshing Golden Files

OpenAlex is a live database — publication counts change daily.
When counts drift, the regression test will fail (correctly).

To update the golden files:

python refresh_golden.py

This regenerates:

  • regression_test_expected_level1.csv
  • regression_test_expected_level2_structure.md

Then re‑run:

python regression_test.py

🧠 How Fallback Institution Matching Works

If OpenAlex cannot find an institution using the org_name field:

  1. The pipeline extracts the domain from WebsiteAddress
  2. It infers a simpler name (e.g., home.dartmouth.edu → “Dartmouth”)
  3. It retries the OpenAlex search

This fixes cases like:

  • “TRUSTEES OF DARTMOUTH COLLEGE”
  • “PRESIDENT AND FELLOWS OF HARVARD COLLEGE”
  • “THE CORPORATION OF BROWN UNIVERSITY”

⚙️ Configuration

At the top of get_counts.py:

LEVEL = 1
RECENT_PUBLICATIONS = 20
MAX_WORKERS = 4

You can change:

  • default level
  • number of recent publications
  • parallelism

Command‑line arguments override the default level.


✔ Recommended Workflow

  1. Edit your input CSV
  2. Run Level 1 or Level 2
  3. When OpenAlex changes and regression tests fail:
    python refresh_golden.py
    python regression_test.py
    
  4. Commit both golden files and the updated code

This keeps your pipeline stable and reproducible.


About

OpenAlex is a bibliographic catalogue of scientific papers, authors and institutions. A philanthropy that endows funds at non-profits with a focus on scholarships for science students uses repository to weigh the impact of an entity. The more published papers that show authors affiliated with that entity, the higher the weight.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages