Alexander Panfilov is a PhD researcher at the ELLIS Institute Tübingen and Max Planck Institute for Intelligent Systems. His work focuses on AI safety, red-teaming large language models, adversarial machine learning, trusted monitors, strategic dishonesty, and the security of LLM-based systems. He previously participated in the MATS 9.0 Google DeepMind stream.
Alexander Panfilov is a PhD researcher at the ELLIS Institute Tübingen and Max Planck Institute for Intelligent Systems. His work focuses on AI safety, red-teaming large language models, adversarial machine learning, trusted monitors, strategic dishonesty, and the security of LLM-based systems. He previously participated in the MATS 9.0 Google DeepMind stream.