Senior AI Research Engineer · Islamabad, Pakistan

I build weather models and production AI systems with measurable results.

At Vayuh.ai and Editable AI, I work on severe-weather models, forecast verification, and source-traceable reports. Earlier, I managed teams building language, image, speech, and vision systems.

Scientific ML Generative systems Computer vision & speech Evaluation & MLOps

Selected projects

Recent work

Weather forecasting is my main research area. My recent independent work also covers agent security and multimodal production infrastructure.

01 / Weather verification Editable AI

Forecast evaluation at continental scale

Demonstrated that Hyper-CONUS v2 achieved lower four-lead mean RMSE than HRRR for 2024 2 m temperature and dew point. At observed ±5 K temperature anomalies, RMSE was 11.2% lower for warm cases and 16.3% lower for cold cases.

Built the matched verification pipeline across 702 initialization times and eight forecasting systems, scoring surface fields against URMA and precipitation against MRMS on native and common grids.

  • Spatiotemporal ML
  • 3 km CONUS grid
  • HPC inference
  • Forecast verification
Read the evaluation
02 / Severe weatherVayuh.ai

Multi-decade severe-weather hindcasts

Built peril-specific U-Net training and hindcast workflows for hail, thunderstorm wind, and tornado risk over a 1980–2025 archive, with spatial, hotspot, ranking, and calibration diagnostics.

Built and audited a 50,000-year, 3 km thunderstorm-wind catalog aligned with hail and tornado simulation years: 155.1 million parent storms, 28.36 billion event-cell rows, and 515 GB of Parquet output.

Validation Developed count, geometry, vector-direction, correlation, calibration, and loss-cost checks before the production catalog was accepted.

Explore OpenSCS
03 / AI securityIndependent R&D

Context-aware prompt-injection detection

Built a cross-encoder pipeline that scores untrusted content together with the agent policy governing how it should be read. At the same 0.70 threshold, ModernBERT outperformed the stored BGE run on Rogue Security 5k and AgentHarm.

0.974Rogue AUC
0.879AgentHarm AUC
0.996S-Labs test AUC

A separate real-data ModernBERT variant reached 0.9955 ROC-AUC and 0.9710 F1 on the 2,101-record S-Labs test split. Those test records had zero exact text matches in the full 440,338-row prepared training set.

Evaluation scope Named stored runs are shown separately; performance varied by dataset. No completed public-baseline run was stored, so this is not a state-of-the-art claim.

Read the case study
04 / Storm reportsEditable AI

Traceable severe-weather reports

Built a deterministic reporting pipeline for hail and tornado cases. It preserves each finding’s source, time, distance, availability state, and limitations.

For an Okeene-area event on 6 May 2024, five weather-data sources returned relevant records. The evidence included 124 SWDI radar-hail signatures, four ground reports, six Storm Events records, observations from four ASOS stations, and MRMS MESH near the property.

  • 0.3 mi nearest report
  • 1.70 in MRMS within 0.6 mi
  • 3.00 in area radar peak
  • 16 / 16 evidence score
  • Address withheld

Interpretation The evidence score reflects source coverage and agreement; it is not a probability. Radar values remain estimates near the property, not direct observations of impact or damage at the structure.

05 / Multimodal systemsActive R&D

Multimodal video production runtime

Built an orchestration layer for short-form episodes across video, dialogue, Foley, lip sync, compositing, and final assembly. Automation of repeated steps and smaller helper models reduced the time to create a three-minute episode by about 40%.

PlanGenerateMeasureRepair

Implementation 107 Python modules, 185 automated tests, resident multi-GPU queues, deterministic manifests, and smallest-unit regeneration.

Professional experience

Research and production engineering since 2019.

Download the full résumé
2022—present

Vayuh.ai

Senior AI Research Engineer

Build peril-specific weather models and catastrophe event sets, including multi-decade hindcasts, correlated multi-peril simulation, calibration, loss-cost analysis, and large-scale production on Perlmutter and Frontier.

  • Spatiotemporal forecasting
  • Distributed training
  • Catastrophe modeling
2025—present

Editable AI

AI Research Lead

Lead regional forecast verification and source-traceable severe-weather reporting, from matched operational benchmarks to property-level evidence packets.

  • Research direction
  • Forecast verification
  • Technical communication
2023—2026

Dashverse

Lead Research Engineer

Managed a team of 5+ and led language, vision, diffusion, and multimodal models from curated data and fine-tuning through production inference.

  • Generative AI
  • Inference systems
  • Research leadership
2021—2023

Pagarba Solutions

Team Lead — AI, MLOps & Optimization

Hired and managed a team of 15+ researchers, MLOps engineers, mobile developers, and software developers. Owned cross-functional ML projects from requirements and technical planning through deployment.

  • Hiring & management
  • Project delivery
  • Mobile inference
2019—2023

WeVoz · STech.ai · Teachinguide

Earlier machine-learning roles

Shipped speech, vision, ranking, forecasting, and classification systems. Results include 4.1% Italian ASR WER, 97.6% RGB anti-spoofing accuracy, and a 70% inference-cost reduction.

  • Speech & vision
  • CUDA & edge inference
  • Cloud deployment

Methods & tools

What I work with

I define the dataset, split, and evaluation protocol before model tuning. For production systems, I add explicit input contracts, failure checks, and reproducible outputs.

I work mainly in Python and PyTorch, with C++ and CUDA when latency or hardware control requires them.

Modeling
PyTorch, Transformers, U-Nets, diffusion, cross-encoders, spatiotemporal models
Systems
Distributed GPU training, Slurm, CUDA, ONNX, Docker, streaming inference
Data
Python, C++, xarray, Dask, Zarr, geospatial and time-series pipelines
Cloud & HPC
NERSC Perlmutter, ORNL Frontier, Azure, GCP

Education

MEng, Machine Learning & Image ProcessingAir University · 2017—2019
Thesis: Sparse Inverse Covariance Estimation
Bachelor of Electrical Engineering (Electronics)Air University · 2013—2017

Availability

Open to senior research engineering roles.

I’m considering roles in scientific ML, applied research, ML systems, and generative media.

Click outside the figure or press Escape to close.