cat ~/research/README.md
My research focuses on AI safety, multimodal learning, and understanding how foundation models encode and process information.
Investigating how audio-language models encode and preserve paralinguistic information across their internal representations.
Demonstrating how auxiliary emotion supervision reshapes affective representation distribution across audio-language model layers.
Novel approach for robust and idempotent voice attribute editing while maintaining speaker identity.
Developing evolving harnesses for computer-use agents to enable safer autonomous systems in dynamic environments.
Novel benchmark for evaluating strategic deception and Theory of Mind in LLM agents through 8,000+ simulated multi-agent games.
Developing methods to make AI systems safer, more controllable, and aligned with human values. Steering mechanisms, safety interventions, and understanding failure modes.
Understanding how audio-language and vision-language models encode and process information across modalities. Investigating information flow and representation learning.
Reverse-engineering neural networks to understand internal representations and computational mechanisms. Using interpretability to improve model behavior and safety.
Designing and analyzing multi-agent systems with focus on strategic interaction, deception, and Theory of Mind. Building evaluation frameworks for agent capabilities.