Careful runs. Trustworthy conclusions.
Mohamed A M Elansary, PhD — candidate for Member of Technical Staff, Evaluation Execution. Measurement discipline under uncertainty, Linux/HPC execution, production agent systems, and honest failure-mode reporting.
Scientific evaluation
- Designed multimodel, multi-basin forecast experiments across hydroclimates.
- Quantified and reduced uncertainty using imperfect USGS, NOAA, and NASA observations.
- Ran reproducible Linux/HPC workflows and Python, R, and Bash data pipelines.
Production delivery
- Builds agentic LLM workflows using GPT, Claude, and Gemini.
- Maintains regression evaluation sets for multi-tenant agent workflows.
- Ships retrieval, routing, isolation, provenance, validation, and monitoring systems.
What I would demonstrate
Execute a small task suite, record scaffold and configuration, inspect traces and results, separate capability failures from task or tooling failures, quantify uncertainty, and deliver a decision-ready memo grounded in the evidence.
Honest fit boundary
I have not published AI-safety research. My research record is environmental engineering and hydrologic forecasting; the relevant transfer is multimodel evaluation, uncertainty quantification, imperfect-data QA, and production agent regression evaluation. I do not claim ownership of a named METR benchmark or prior dangerous-capability evaluation.
Role and logistics
Berkeley, CA; technical team members are in the office 3–5 days per week, with flexibility. Berkeley hybrid is a relocation-with-package ask.
Posted compensation: $328,380–$578,583 per year, with junior/mid-level and senior subranges described separately.
Official role: METR / Lever