Portrait of Essa Jan

Essa Jan

2nd Year Masters @ Brown University

I am a second year MS CS student at Brown University, advised by Stephen Bach. In my previous, truly amazing life, I completed my undergrad in CS at LUMS, and was advised by Professor Fareed Zaffar, Professor Yasir Zaki, and Faizan Ahmed. I am currently interning at Preference Model in the post training team, where I am working on agent trajectory correction, error recovery, reasoning trace generation, and RL environments to improve LLM post-training capabilities.

The next generation of AI systems will not be defined only by their pretrained knowledge, but by their ability to learn, reason, and improve throughout their lifecycle. My research interests lie in post-training approaches that enable models to improve their capabilities, adapt over time, and remain safe. I am particularly interested in RL, self-improvement, and developing mechanisms to evaluate and enhance model capabilities while ensuring alignment.

News

June 2026
Joined Preference Model as RL Engineering Intern
August 2025
Our recent work on Evaluating the Transferability of Injected Knowledge in LLMs made it into the EMNLP Findings track.
December 2024
Our paper on MultitaskBench was accepted at COLING 25, focusing on safety gaps in LLM fine-tuning.

Publications

MultitaskBench: Unveiling and Mitigating Safety Gaps in LLMs Fine-tuning

Essa Jan, Nouar Aldahoul, Moiz Ali, Faizan Ahmad, Fareed Zaffar, Yasir Zaki

2025 · COLING 25

Investigated how fine-tuning on downstream tasks affects the safety guardrails of large language models. We develop a comprehensive benchmark to evaluate safety degradation and the robustness of different safety solutions across multiple task domains.

Understanding or Imitation? Auditing Conceptual Understanding and Reasoning in Large Language Models

Saram Hassan, Essa Jan, Ramneet Kaur, Eric Yeh, Fareed Zaffar, Ashish Gehani

2026 · ACM REP

We reproduced and extended the "Potemkin Understanding in LLMs" paper, which claims that current benchmarks do not correctly evaluate LLM understanding and conceptual representation. We reran both evaluation techniques mentioned in the paper and extended them to smaller, reasoning, and quantized models. The directional claims hold, but reported scores are unstable, varying up to 31% across identical runs. Reasoning models reduce incoherence sharply, suggesting these failures depend on the inference regime, not the model class.

Teaching