Bengaluru, India · Senior AI Architect @ BLUE

Shubham
Baid

$> _

AI Systems Engineer taking video and multimodal AI from research to production. Specializing in perception-aware compression, VLM edge pipelines, and real-time inference optimization under strict resource budgets.

0+
Years Exp.
0+
Cameras Deployed
0%
Max Bitrate Cut
2 Pubs
1 Patent App
Shubham Baid
Shubham Baid AI Systems Engineer shubhambaid99@gmail.com
HackZurich Winner NASA Space Apps Winner
SCROLL TO EXPLORE
PyTorch/ Vision-Language Models (LLaVA, Qwen-VL, InternVL)/ TensorRT & CUDA/ Perception-Aware Compression/ NVIDIA DeepStream/ GStreamer & RTSP/ Model Quantization (INT8/FP16)/ YOLO/ OpenCV & Optical Flow/ Multi-Agent Systems (CrewAI, LangGraph)/ NVIDIA Jetson & ARM/NEON/ Docker & Linux/ Python & C++/ PyTorch/ Vision-Language Models (LLaVA, Qwen-VL, InternVL)/ TensorRT & CUDA/ Perception-Aware Compression/ NVIDIA DeepStream/ GStreamer & RTSP/ Model Quantization (INT8/FP16)/ YOLO/ OpenCV & Optical Flow/ Multi-Agent Systems (CrewAI, LangGraph)/ NVIDIA Jetson & ARM/NEON/ Docker & Linux/ Python & C++/

I am an AI Systems Engineer with 5+ years of experience taking video and multimodal AI systems from research to production, shipped to 1,000+ cameras across schools, banks, warehouses, and moving metro trains.

I specialize in solving AI problems under severe real-world constraints — limited network bandwidth, heterogeneous edge hardware, and strict latency budgets — by questioning conventional assumptions and redesigning systems around deployment reality rather than reaching for more compute.

At BLUE, this thinking produced a perception-aware video compression system for machine vision that cuts bitrate by a 40% median (up to 69%) beyond H.265 with under 1% impact on downstream detection accuracy — now running across 1,000+ production cameras and resulting in an arXiv publication and patent filing.

Vision-Language Models Perception Compression TensorRT & CUDA GStreamer & DeepStream Edge AI & NEON Agentic Workflows Python & C++
Engineering Snapshot
Primary Expertise Video & Multimodal AI
Cameras Deployed 1,000+ Production
Bandwidth Saved 40% - 69% Bitrate Cut
Publications 2 (arXiv, TEST Eng)
Patents 1 Filed
Global Hackathons Winner (Zurich, NASA)
Location Bengaluru, India

Senior AI Architect

BLUE (myblue.ai) · Bengaluru, India
Apr 2025 – Present
  • Questioned the default assumption that surveillance video should be compressed for human viewing, and architected a perception-aware compression pipeline tuned for downstream AI instead: 40% median (up to 69%) bitrate reduction beyond H.265, with detection precision/recall deltas under 1% and VLM output deviation under 0.1%.
  • Iteratively redesigned the compression pipeline over several months through large-scale benchmarking against H.264/H.265, using custom semantic metrics to guide architectural decisions rather than perceptual quality alone.
  • Developed Foreground VMAF, an evaluation methodology for benchmarking perception-aware compression by focusing on semantically relevant foreground regions rather than static background — now the team's benchmarking standard, and the basis for an arXiv publication and patent filing.
  • Architected the complete edge AI pipeline from RTSP ingest through perception-aware compression, streaming, decoding, and VLM inference, deployed across 1,000+ cameras in offices, banks, warehouses, and moving metro trains while preserving operational reliability under constrained network conditions.
  • Established a benchmarking framework evaluating Vision-Language Models across latency, throughput-per-watt, hallucination rate, and downstream task accuracy to guide production deployment on constrained edge hardware.
  • Designed a Vision-Language Model pipeline for automated restaurant order validation from CCTV, replacing manual audits by reconciling live observations with POS transactions through structured reasoning and fuzzy matching.
Perception-Aware CompressionOptical FlowVLMs (LLaVA, Qwen-VL)RTSP StreamingEdge AIForeground VMAF

AI Engineer

CAFU · Dubai, UAE (Remote)
Jan 2025 – Apr 2025
  • Re-architected CAFU's ETA prediction pipeline by introducing real-time behavioral signals, reducing delayed fulfillment from 13% to 8% and cutting severe SLA breaches by 42%.
  • Engineered an agentic LLM workflow for marketing content generation, lifting CTR by 15% and cutting copywriting turnaround time by 50%.
  • Designed autonomous B2B lead acquisition using multi-modal LLM agents, processing 1,400 prospects per month and cutting manual prospecting hours by 90%.
Agentic WorkflowsMulti-Modal LLMsCrewAI / LangGraphReal-Time AnalyticsPython

Senior AI Engineer

Avathon (formerly SparkCognition) · Bengaluru, India
May 2022 – Jan 2025
  • Led AI for the VAIA School Safety Suite, real-time video understanding across 100+ cameras in live, safety-critical US school deployments owning model deployment, pipeline reliability, and inference optimization.
  • Challenged the default reliance on fixed-class object detection for context-rich scenes and integrated Vision-Language Models into classical CV pipelines for complex scene understanding — an early production adoption of VLMs ahead of broader industry uptake.
  • Designed and built a gamified active-learning annotation tool to break the manual-labeling bottleneck slowing dataset curation, accelerating continuous retraining loops for production CV models.
  • Trained and shipped detection and classification models serving thousands of production cameras across enterprise and industrial sites over my tenure.
  • Cut inference CPU usage by 50% at sustained real-time throughput via ARM/NEON-optimized pipelines with INT8 quantization.
Multi-Camera CVTensorRTARM / NEONINT8 QuantizationActive LearningVision-Language Models

Senior AI Engineer

Integration Wizards (acquired by SparkCognition) · Bengaluru, India
Sep 2020 – May 2022
  • Owned a production ALPR system (YOLO detection + OCR) end to end, with hybrid CPU/GPU inference optimized for real-time video streams on a minimal resource footprint.
  • Migrated legacy CV models onto NVIDIA DeepStream and TensorRT GPU pipelines to unlock real-time throughput, establishing the team's production video inference architecture.
  • Delivered PPE detection, fall-arrester detection, and vehicle classification models from dataset curation through deployment — this core CV technology was central to the company's acquisition by SparkCognition.
NVIDIA DeepStreamTensorRTYOLOALPR & OCRHybrid CPU/GPU Inference

AI Systems & Multimodal

Foundational and modern deep learning models, multimodal reasoning, quantization, and agentic workflows.

PyTorch Vision-Language Models LLaVA Qwen-VL InternVL INT8 / FP16 Quantization Model Distillation RAG Multi-Agent Systems (CrewAI, LangGraph)

Computer Vision & Video

Real-time video ingest, optical flow, decoding pipelines, and perception-aware compression architecture.

YOLO Family OpenCV Optical Flow GStreamer NVIDIA DeepStream RTSP / SRT Streaming FFmpeg Hardware Encode / Decode

Performance & Edge

Maximizing throughput-per-watt and latency budgets on constrained GPU and ARM edge targets.

TensorRT CUDA NVIDIA Jetson ARM / NEON Optimization Real-time Pipeline Design Latency & Throughput Benchmarking

Programming & Infrastructure

Robust production backends, containerized microservices, and system programming languages.

Python C++ Docker Linux Systems FastAPI Vector Databases Git & CI/CD
Production & Research
2025 - Present

Perception-Aware Video Compression System

Architected a groundbreaking compression pipeline tuned for downstream machine vision rather than human eyeballs. Achieves 40% median (up to 69%) bitrate reduction beyond H.265 with <1% downstream detection impact across 1,000+ cameras on moving trains, banks, and warehouses.

Optical FlowPerception CompressionRTSPVLMsForeground VMAF
Production
Avathon

VAIA School Safety Suite

Led AI engineering for a safety-critical real-time video understanding suite deployed across 100+ live cameras in US schools. Built multi-camera object detection, person re-ID, and ARM/NEON optimized pipelines.

Multi-Camera CVTensorRTARM / NEONINT8
Production
BLUE

QSR Vision Order Validation Pipeline

Designed an automated restaurant order validation pipeline using Vision-Language Models on CCTV footage, reconciling physical tray items with POS transactions via structured reasoning and fuzzy matching.

VLMStructured ReasoningLive CCTVAutomation
HackZurich 2021 Winner
Europe's Largest Hackathon

YetiCoach — Real-Time Sports Computer Vision

Built a real-time ski technique coaching system from action camera footage for Sunrise, Huawei, and Swiss-Ski. Analyzes ski angles, posture, and technique on the edge; officially recognized by the Swiss-Ski federation.

Sports CVReal-Time Pose AnalysisEdge Processing
NASA Space Apps 2021 Winner
Global Winner

Project Aegir — Satellite & UAV Marine AI

Computer vision platform leveraging satellite and UAV imagery to detect, classify, and predict oceanic debris drift patterns for automated cleanup coordination.

Satellite CVUAV ImageryDebris Tracking
Production
CAFU

Agentic LLM Prospecting & Content Generation

Architected autonomous B2B lead acquisition agents processing 1,400 prospects/month (cutting manual effort by 90%) and agentic marketing workflows boosting CTR by 15%.

Agentic AIMulti-Modal LLMsCrewAIAutomation
Production
Integration Wizards

High-Throughput ALPR System

End-to-end Automatic License Plate Recognition system (YOLO detection + OCR) with hybrid CPU/GPU inference optimized for high-fps video streams on minimal hardware.

YOLOOCRDeepStreamTensorRT
Production Tooling
Avathon

Gamified Active-Learning Annotation Tool

Designed and built an active-learning annotation environment to break manual labeling bottlenecks, drastically speeding up dataset curation and continuous model retraining loops.

Active LearningDataset CurationRetraining Loops
arXiv, 2026 · Machine Vision & Compression

BLUE: A Stale-Pixel Optical-Flow Compositor for Entropy-Efficient Surveillance Video Encoding

Presents an optical-flow based compositor designed specifically for perception-aware surveillance encoding, removing temporal redundancy for machine vision consumers without sacrificing downstream detection accuracy.

Perception CompressionOptical FlowSurveillance VideoarXiv
Patent Application · Intellectual Property

Perception-Aware Surveillance Video Compression for Machine Vision

Patent application covering the architecture, foreground-aware masking, and evaluation methodologies for compressing video streams optimized for AI detectors and VLMs.

Patent Filed
TEST Engineering and Management · May 2020

Detection of Different Degrees of Skin Burn using YOLOv3

Applied object detection models for automated severity classification of skin burns from medical imagery.

Computer VisionYOLOv3Medical Imaging

Education

B.Tech, Computer Science

REVA University, Bengaluru · 2017 – 2021
  • Best Outgoing Student award recipient.
  • Founded GDSC REVA (Google Developer Student Club).

Awards & Honors

🏆 Winner
HackZurich 2021

Europe's largest hackathon. Built YetiCoach (real-time sports CV for Sunrise, Huawei, Swiss-Ski).

🚀 Winner
NASA Space Apps 2021

Global winner with Project Aegir (satellite and UAV marine debris detection AI).

Specialized Certifications

NVIDIA Jetson AI Specialist
Intel Edge AI Specialist
TensorFlow Developer Certificate (Google)

I am always interested in discussing technical challenges across AI systems architecture, edge video analytics, perception compression, and bringing frontier multimodal models to real-world deployment.

Let's Build Together

Based in Bengaluru, India · Open to technical leadership, consulting, and global collaboration.