A comprehensive data processing system that automatically extracts, transforms, and loads data from unstructured files, generates schemas, tracks schema evolution, and enables natural language querying.
- Overview
- Directory Structure
- System Architecture
- Code Flow
- Module Details
- Usage Examples
- API Endpoints
The Sorting Hat is an intelligent ETL (Extract, Transform, Load) pipeline that:
- Parses unstructured files (HTML, JSON, CSV, etc.) and detects data formats
- Extracts entities using Named Entity Recognition (NER)
- Generates database schemas automatically
- Tracks schema evolution over time
- Translates natural language queries to SQL
- Multi-format Detection: JSON, CSV, HTML tables, YAML, Key-Value pairs, SQL, and more
- NER Integration: SpaCy and GLiNER for entity extraction
- Automatic Schema Generation: PostgreSQL and MongoDB schemas
- Schema Evolution Tracking: Version control for schemas with migration scripts
- Natural Language Queries: Convert plain English to SQL queries
TheSortingHat/
β
βββ Core Python Files
β βββ backend.py # FastAPI REST API server
β βββ etl_parser.py # File parsing and format detection
β βββ schema_generator.py # Schema generation from parsed data
β βββ schema_evolution.py # Schema versioning and evolution tracking
β βββ query_translator.py # Natural language to SQL translation
β βββ main.py # Standalone script for testing
β
βββ Data Directories (auto-created)
β βββ uploads/ # Uploaded files storage
β βββ schemas/ # Generated schema files (SQL/JSON)
β βββ records/ # Query results and ETL summaries
β βββ schema_registry/ # Schema version registry (JSON)
β
βββ Input Files (examples)
βββ input.txt
βββ input2.html
βββ input3.txt
βββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ
β User/Client Request β
ββββββββββββββββββββββββ¬βββββββββββββββββββββββββββββββββββββββ
β
βΌ
βββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ
β backend.py (FastAPI) β
β ββββββββββββββββ ββββββββββββββββ ββββββββββββββββ β
β β POST /upload β β GET /schema β β POST /query β β
β ββββββββ¬ββββββββ ββββββββ¬ββββββββ ββββββββ¬ββββββββ β
βββββββββββΌβββββββββββββββββββΌβββββββββββββββββββΌββββββββββββββ
β β β
βΌ βΌ βΌ
βββββββββββββββββββ βββββββββββββββββββ βββββββββββββββββββ
β etl_parser.py β βschema_evolution β βquery_translator β
β β β .py β β .py β
β β’ Parse file β β β β β
β β’ Detect β β β’ Track β β β’ Translate β
β formats β β versions β β NL β SQL β
β β’ Extract β β β’ Generate β β β’ Execute β
β entities β β migrations β β queries β
ββββββββββ¬βββββββββ βββββββββββββββββββ βββββββββββββββββββ
β
βΌ
βββββββββββββββββββ
βschema_generator β
β .py β
β β
β β’ Infer types β
β β’ Generate β
β schemas β
β β’ Export DDL β
βββββββββββββββββββ
1. FILE UPLOAD
ββ> backend.py: POST /upload
β
ββ> Save file to uploads/
β
ββ> Parse file content
β
βΌ
2. ETL PARSING (etl_parser.py)
β
ββ> ETLFragmentDetector.run_all()
β ββ> Detect JSON, CSV, HTML, YAML, etc.
β ββ> Extract fragments with confidence scores
β ββ> Mark parent-child relationships
β
ββ> NEREnricher.enrich_all_fragments()
β ββ> SpaCy: Standard entity recognition
β ββ> GLiNER: Domain-specific entities (product_id, price, etc.)
β ββ> Pattern matching: Emails, URLs, product IDs
β
ββ> Normalizer.normalize()
β ββ> Convert fragments to structured data (dicts/lists)
β
ββ> group_fragments_by_entity()
β ββ> Group fragments by product_id or entity
β
ββ> merge_fragments_for_entity()
ββ> Merge data from multiple fragments
β
βΌ
3. SCHEMA GENERATION (schema_generator.py)
β
ββ> SchemaBuilder.build_all_schemas()
β ββ> TypeInferrer.infer_type()
β β ββ> Infer data types (string, integer, date, etc.)
β β
β ββ> Detect primary keys and indexes
β β
β ββ> Categorize entities (products, documents, etc.)
β
ββ> SchemaExporter.export_all()
ββ> Generate PostgreSQL DDL
ββ> Generate MongoDB JSON Schema
β
βΌ
4. SCHEMA EVOLUTION (schema_evolution.py)
β
ββ> SchemaEvolutionManager.process_new_schema()
β β
β ββ> SchemaRegistry.register_schema()
β β ββ> Save to schema_registry/registry.json
β β
β ββ> SchemaDiffer.diff_schemas()
β β ββ> Compare with previous version
β β ββ> Detect: added/removed/renamed fields
β β ββ> Detect: type changes, nullability changes
β β
β ββ> MigrationGenerator.generate_migration()
β ββ> Generate forward migration SQL
β ββ> Generate rollback SQL
β ββ> Generate compatibility views
β
ββ> Return evolution results
β
βΌ
5. RESPONSE
ββ> Return JSON with:
ββ> Processing summary
ββ> Fragments detected
ββ> Entities found
ββ> Schemas generated
ββ> Schema evolution info
1. NATURAL LANGUAGE QUERY
ββ> backend.py: POST /query
β
ββ> NaturalLanguageQuerySystem.query()
β
ββ> Load schemas from registry
β
ββ> LLMQueryTranslator.translate_to_sql()
β ββ> Use Ollama (phi3:mini) to convert NL β SQL
β
ββ> QueryValidator.validate_sql()
β ββ> Check for dangerous operations
β ββ> Validate column names
β
ββ> QueryExecutor.execute_sql()
ββ> Execute query (mock or real DB)
β
βΌ
2. RESPONSE
ββ> Return:
ββ> Translated SQL query
ββ> Query results
ββ> Execution time
ββ> Query ID for later retrieval
Purpose: Parse unstructured files and detect data formats
Key Classes:
ETLFragmentDetector: Detects different data formats in textNEREnricher: Adds Named Entity Recognition to fragmentsNormalizer: Converts fragments to structured dataDetectedBlock: Data class representing a detected fragment
Format Detection Priority:
- JSON-LD (highest priority)
- JSON
- YAML Frontmatter
- HTML Tables
- CSV
- Key-Value pairs
- JavaScript Objects
- SQL
- Raw Text (lowest priority)
Output:
{
'fragments': [DetectedBlock, ...],
'summary': {'JSON': 5, 'CSV': 2, ...},
'records': [...],
'entity_index': {...},
'grouped_entities': {...},
'merged_entities': {...}
}Key Functions:
parse_file(text, enable_ner=True): Main entry point
Purpose: Generate database schemas from parsed entity data
Key Classes:
TypeInferrer: Infers data types from valuesSchemaBuilder: Builds schemas from merged entitiesSchemaExporter: Exports schemas to PostgreSQL/MongoDB
Type Inference:
- Detects: integers, decimals, strings, dates, booleans, arrays, objects
- Special handling: emails, URLs, prices, product IDs
- Confidence scoring for type inference
Output:
- PostgreSQL DDL files:
postgresql_<category>.sql - MongoDB JSON Schema files:
mongodb_<category>.json
Key Functions:
generate_schemas_from_etl_result(etl_result, output_dir): Main entry point
Purpose: Track schema changes over time and generate migrations
Key Classes:
SchemaRegistry: Central registry for all schema versionsSchemaDiffer: Compares schemas and detects changesMigrationGenerator: Generates SQL migration scriptsSchemaEvolutionManager: Main interface for schema evolution
Change Detection:
- Field added/removed
- Field renamed (heuristic-based)
- Type changes
- Nullability changes
- Breaking vs non-breaking changes
Output:
schema_registry/registry.json: Version history- Migration SQL files (forward, rollback, views)
Key Functions:
SchemaEvolutionManager.process_new_schema(): Process new schema versionSchemaEvolutionManager.get_schema_history(): Get version history
Purpose: Translate natural language queries to SQL
Key Classes:
LLMQueryTranslator: Uses Ollama LLM to translate queriesQueryValidator: Validates generated SQL queriesQueryExecutor: Executes queries (mock or real DB)NaturalLanguageQuerySystem: Complete NL query system
Translation Process:
- Load schema from registry
- Build schema context for LLM
- Generate SQL using LLM (Ollama phi3:mini)
- Validate SQL query
- Execute query
Key Functions:
NaturalLanguageQuerySystem.query(): Main entry point
Purpose: REST API server that orchestrates all modules
Endpoints:
| Method | Endpoint | Description |
|---|---|---|
POST |
/upload |
Upload and process file through ETL pipeline |
GET |
/schema |
Get schema(s) by source_id or entity_type |
GET |
/schema/history |
Get schema evolution history |
POST |
/query |
Execute natural language query |
GET |
/records |
Get query results or source summaries |
GET |
/entities |
List all available entity types |
GET |
/health |
Health check endpoint |
GET |
/ |
API information |
Key Features:
- File upload with size validation
- CORS enabled
- Error handling
- JSON responses
Purpose: Example script showing how to use the modules directly
Usage:
python main.pyWhat it does:
- Reads
input2.html - Parses with NER enabled
- Generates schemas
- Tracks schema evolution
- Prints detailed analysis
# Start the server
python backend.py
# Upload a file
curl -X POST "http://localhost:5000/upload" \
-F "file=@input2.html" \
-F "source_id=test-001"
# Get schema
curl "http://localhost:5000/schema?entity_type=products"
# Query with natural language
curl -X POST "http://localhost:5000/query" \
-H "Content-Type: application/json" \
-d '{"query": "Show me all products with price less than 50"}'from etl_parser import parse_file
from schema_generator import generate_schemas_from_etl_result
from schema_evolution import SchemaEvolutionManager
# Parse file
with open("input2.html", "r") as f:
text = f.read()
result = parse_file(text, enable_ner=True)
# Generate schemas
schemas = generate_schemas_from_etl_result(result, output_dir='./schemas')
# Track evolution
schema_manager = SchemaEvolutionManager('./schema_registry')
for entity_id, schema in schemas.items():
entity_type = schema.get('category', entity_id)
evolution_result = schema_manager.process_new_schema(entity_type, schema)
print(f"Schema version: {evolution_result['version'].version}")from query_translator import NaturalLanguageQuerySystem
query_system = NaturalLanguageQuerySystem('./schema_registry')
# Query
result = query_system.query(
"Show me all products with price less than 50",
entity_type="products",
database_type="postgresql"
)
print(result['query']) # Generated SQL
print(result['results']) # Query resultsUpload and process a file through the ETL pipeline.
Request:
file: File to upload (multipart/form-data)source_id: Optional source identifierversion: Optional version stringenable_ner: Enable NER (default: true)
Response:
{
"success": true,
"data": {
"source_id": "...",
"filename": "...",
"file_size": 12345
},
"processing": {
"fragments_detected": 10,
"entities_found": 5,
"schemas_generated": 3
}
}Get schema(s) by source_id or entity_type.
Query Parameters:
source_id: Optional source IDentity_type: Optional entity type
Response:
{
"success": true,
"entity_type": "products",
"version": 1,
"schema": {...}
}Execute a natural language query.
Request Body:
{
"query": "Show me all products with price less than 50",
"entity_type": "products",
"database_type": "postgresql"
}Response:
{
"success": true,
"query_id": "...",
"translated_query": "SELECT * FROM products WHERE price < 50",
"results": [...]
}# Install Python dependencies
pip install fastapi uvicorn beautifulsoup4 python-multipart
# Optional: For NER features
pip install spacy gliner
python -m spacy download en_core_web_sm
# Optional: For query translation
# Install Ollama and pull model
ollama pull phi3:mini# Start the API server
python backend.py
# Or using uvicorn directly
uvicorn backend:app --host 0.0.0.0 --port 5000
# Or using FastAPI CLI
fastapi run backend.py --port 5000- API Documentation: http://localhost:5000/docs
- Alternative Docs: http://localhost:5000/redoc
- Health Check: http://localhost:5000/health
Input File
β
[etl_parser.py] β Fragments + Entities
β
[schema_generator.py] β Database Schemas
β
[schema_evolution.py] β Versioned Schemas
β
[query_translator.py] β Natural Language Queries
β
Query Results
Detected data blocks in the input file (JSON objects, CSV rows, HTML tables, etc.)
Named entities extracted from fragments (product IDs, prices, emails, etc.)
Multiple fragments grouped together by entity ID (e.g., all fragments mentioning "prod-1001")
Tracking changes to schemas over time, generating migrations, and maintaining backward compatibility
Converting plain English questions into SQL queries using LLM
- NER models (SpaCy, GLiNER) are optional - the system works without them, but works best with them.
- Query translation requires Ollama with phi3:mini model (or any model you can manually configure)[optional]
- Schema registry is stored in JSON format in
schema_registry/ - All generated files are saved in respective directories
This is a comprehensive ETL pipeline system. Each module can be used independently or together through the FastAPI backend.