lvlm
Here are 31 public repositories matching this topic...
🔥🔥🔥 A curated list of papers on LLMs-based multimodal generation (image, video, 3D and audio).
-
Updated
Apr 4, 2025 - HTML
OpenThinkIMG is an end-to-end open-source framework that empowers LVLMs to think with images.
-
Updated
Jun 1, 2025 - Jupyter Notebook
up-to-date curated list of state-of-the-art Large vision language models hallucinations research work, papers & resources
-
Updated
Aug 28, 2026
[NeurIPS 2024] This repo contains evaluation code for the paper "Are We on the Right Way for Evaluating Large Vision-Language Models"
-
Updated
Sep 26, 2024 - Python
Awesome-Backdoor-on-LMMs is a collection of state-of-the-art, novel, exciting backdoor methods on LMMs (VLPs, TDMs, VLMs, and Agents).
-
Updated
Aug 31, 2026
Latest Advances on (RL based) Multimodal Reasoning and Generation in Multimodal LLMs
-
Updated
Aug 22, 2026
[ICCV'25] The official code of paper "Combining Similarity and Importance for Video Token Reduction on Large Visual Language Models"
-
Updated
Jan 13, 2026 - Python
📜 Paper list on decoding methods for LLMs and LVLMs
-
Updated
Aug 19, 2026
[ICCV 2025] HQ-CLIP: Leveraging Large Vision-Language Models to Create High-Quality Image-Text Datasets
-
Updated
Aug 6, 2025
[AAAI 2025] HiRED strategically drops visual tokens in the image encoding stage to improve inference efficiency for High-Resolution Vision-Language Models (e.g., LLaVA-Next) under a fixed token budget.
-
Updated
Apr 18, 2025 - Python
CLIP-MoE: Mixture of Experts for CLIP
-
Updated
Oct 10, 2024 - Python
[ICLR 2026] AgilePruner: An Empirical Study of Attention and Diversity for Adaptive Visual Token Pruning in Large Vision-Language Models
-
Updated
Aug 11, 2026 - Python
The official repository of "SmartAgent: Chain-of-User-Thought for Embodied Personalized Agent in Cyber World".
-
Updated
Jul 27, 2026 - Python
Code for ICLR 2025 Paper: Visual Description Grounding Reduces Hallucinations and Boosts Reasoning in LVLMs
-
Updated
May 7, 2025 - Python
LEMMA: An effective and explainable way to detect multimodal misinformation with LVLM and external knowledge augmentation, incorporating the intuition and reasoning capbility inside LVLM.
-
Updated
Jun 4, 2025 - Jupyter Notebook
[CVPR'25] Antidote: A Unified Framework for Mitigating LVLM Hallucinations in Counterfactual Presupposition and Object Perception
-
Updated
Oct 11, 2025 - Python
A benchmark dataset and simple code examples for measuring the perception and reasoning of multi-sensor Vision Language models.
-
Updated
Dec 27, 2024 - Python
Add this topic to your repo
To associate your repository with the lvlm topic, visit your repo's landing page and select "manage topics."