I'm an AI researcher working on Trustworthy AI, Large Language Models, and AI Agents.
Research Interests: LLM Safety · Privacy · Representation Steering · Decoding · RAG · AI Agents
I'm an AI researcher working on Trustworthy AI, Large Language Models, and AI Agents.
Research Interests: LLM Safety · Privacy · Representation Steering · Decoding · RAG · AI Agents
[COLM 2026] Do LLMs Know What Is Private Internally? Probing and Steering Contextual Privacy Norms in Large Language Model Representations
📜 Paper list on decoding methods for LLMs and LVLMs
[COLING 2025] Piecing It All Together: Verifying Multi-Hop Multimodal Claims.
Python 10
[CIKM 2024] Trojan Activation Attack: Attack Large Language Models using Activation Steering for Safety-Alignment.