I’m Ruchira, a PhD Researcher in NLP and AI at the Department of Computer Science, University of Copenhagen, supervised by Dr. Anders Søgaard. I’m affiliated with the CoAStaL Lab and the Centre for Philosophy of AI. Currently visiting the LRE Lab at ETH Zurich, and interning at Refractal AI, an agentic security startup in London.
I’m on the job market from early next year. I’m open to research and research engineering roles in evaluation, safety, and security, as well as to technical policy and governance roles working on evaluation standards, auditing, and transparency. If you have an interesting role in mind, do get in touch.
Research
I work on evaluation methodology. The questions that interest me are ones like: how do evaluation results change when we change the evaluation setup? How well do efficient evaluation methods work across different types of benchmark data? How can evaluation reporting be improved for transparency and governance purposes?
Some of the work behind those questions:
- Evaluating Adjective-Noun Compositionality in LLMs: Functional vs Representational Perspectives — the same capability question gives you different answers depending on whether you test behaviour or inspect representations (*SEM, ACL 2026).
- Does Item Response Theory Generalize Across Languages? Evidence from Cross-Lingual LLM Benchmarks — whether psychometric methods for efficient evaluation still hold when the benchmark data changes language (forthcoming in EMNLP Proceedings, 2026).
- EvalCards: A Framework for Standardized Evaluation Reporting — a reporting format that makes evaluation results auditable and interpretable for research, industry, and policy audiences, rather than merely comparable between models.
At my internship, I’m also looking into safety and security evaluation for agentic systems: building an evaluation harness for AI agents, and working on privacy evaluations, prompt injection, and red teaming under real regulatory requirements.
Extracurriculars
- Working Member, EvalEval Coalition (2025–Present). Contributing to community efforts around robust and responsible AI evaluation.
- Venture Fellow, Creator Fund (2025–Present). Technical diligence on product and infrastructure for early-stage AI startups.
- Open Science Lead, Cohere Labs (2024–Present). Driving research engagement across a global community of 200+ researchers.
Teaching
Teaching Assistant, NLP Course, University of Copenhagen (2023–Present). Two years in, it remains one of the most rewarding parts of my PhD. There’s something genuinely satisfying about helping students get their first foothold in the field.
Contact
Whether you want to debate evaluation metrics, talk AI safety and policy, or just say hi, find me on LinkedIn or drop a line at this email.
Based in Copenhagen. Coffee recommendations always welcome.
