Next Arc is a content-based anime recommendation system that helps users discover new anime based on what they already enjoy. The system analyzes anime metadata such as genres, tags, and themes to identify similarities between titles and generate meaningful recommendations.
The project applies TF-IDF vectorization and cosine similarity to model content relationships and is delivered through an interactive Streamlit web interface, making anime discovery intuitive, fast, and engaging.
-
Data Collection
The system uses an anime dataset containing metadata like titles, genres, tags, and themes. -
Text Processing & Feature Extraction
Metadata fields are combined and transformed into numerical vectors using TF-IDF (Term Frequency–Inverse Document Frequency) to capture the importance of each term across all anime. -
Similarity Calculation
Cosine similarity is computed between anime vectors to measure how closely related each title is to the others. -
Recommendation Generation
Given a selected anime, the system ranks all other titles by similarity scores and returns the top matches as personalized recommendations.
- Programming Language: Python
- Libraries:
pandas,numpy— data manipulationscikit-learn— TF-IDF vectorization & cosine similaritystreamlit— interactive web interfacepickle— saving and loading preprocessed data
- Frontend: Streamlit for quick, responsive UI
- Data: Kaggle Anime Dataset (metadata including genres, tags, themes)
https://www.kaggle.com/datasets/dbdmobile/myanimelist-dataset
- Clone the repository
git clone https://github.com/aryandas2911/Next-Arc
cd next-arc- Install dependencies
pip install -r requirements.txt- Run the Streamlit app
streamlit run app.py- Interact with the app
- Select an anime you’ve watched
- Get a ranked list of recommended anime based on content similarity
- Explore new anime with their metadata
- Dynamic Data Updates: Automatically fetch new anime and update the dataset without manual preprocessing.
- Enhanced Features: Include more metadata such as ratings, studios, and user reviews to improve recommendation quality.
- Hybrid Recommendation: Combine content-based filtering with collaborative filtering for personalized recommendations.
- Deployment Enhancements: Containerize the app using Docker or deploy on cloud platforms like Render or Heroku for better scalability.
- Performance Optimization: Use sparse matrices or approximate nearest neighbor search to handle larger datasets efficiently.