t=1000 · noise

I make generative models fast enough to ship.

ML Engineer at fal, where generative media runs in production. Inference optimization, on-device AI, and the systems large enterprises bet on.

-- fps · -- ms/frame · -- particles

t≈700 · experience

A model that works in a notebook is an argument. A model that holds its latency target under real traffic is a product. My job is turning the first into the second.

fal

ml engineer · 03/2025 – present · remote

E2E

enterprise systems, built and owned end to end

Adobe. Shopify. Quora. Perplexity. Production inference systems from research to deployment, run against each company's latency, throughput, and reliability targets, with one owner for the whole system: me.

SRE

production clients on my pager

Canva, Netflix, Epic Games, Artlist, Freepik, World Labs. When their systems break, I'm in Grafana and Datadog tracing the incident and shipping the fix.

MCP

AI automations for support and operations

Designed MCP-based automations for support and operations workflows, cutting manual effort on repetitive tasks.

Styldod

ml intern · 08/2024 – 10/2024 · remote

2.5×

faster text-to-image and image-to-image

DeepCache and OneDiff, layered into production diffusion pipelines. Same models, same outputs, less than half the latency.

16k+

images collected into a training dataset

An automated collection pipeline built on Scrapy, BeautifulSoup, and Selenium, with scraping and quality filtering built in.

t≈500 · selected work

01

HedgeMind

Autonomous agents running a portfolio on live market data.

Built for Inter IIT Tech Meet 14.0, on the Pathway problem statement, with the Cynaptics Club team. Kafka streams feed live prices, Reddit and Twitter sentiment, and FRED macro data into Pathway-based agents that reason over the stream and act. SARIMAX and Chronos handle forecasting, a custom MCP layer wires agents to models, and a Django + Next.js dashboard tracks P&L and anomaly alerts in real time. Not a backtest notebook: a streaming system that holds up while the market moves.

Pathway · Kafka · agents + MCP · SARIMAX / Chronos · Inter IIT 14.0

github.com/CynapticsAI/Pathway_InterIIT_14demo video

02

Curio

Crawl any website and talk to it.

The language model runs entirely in your browser over WebGPU (WebLLM), so nothing you ask ever leaves your device. Not as a policy promise, as an architecture: there is no server to send it to. Embedding it on a site takes one line of code. It is also a working demo of the thing I do at fal: inference that runs where the user is.

WebGPU · WebLLM · on-device · one-line embed

curio-umber-five.vercel.app

live widget · pending embed

03

LLM Distillation & Pruning

Knowledge distillation for models the size of Llama 3.1 70B, on hardware that shouldn't allow it.

Two stages. First, run the teacher once and cache its logits. Then train the student against the cached distributions on KL divergence, so the teacher never has to share the GPU with the student. VRAM stays low; performance stays put. Built at Cynaptics Club, IIT Indore.

Llama 3.1 70B scale · logit caching · KL divergence · low VRAM

04

BossForge

Photograph your messy desk. Fight it.

Your mess, your monster: an AI pipeline catalogs the clutter in the photo, designs a multi-phase boss out of it, renders it as a 3D model, and drops you into an arena against your own coffee mug. Image understanding, 3D generation, and a game stacked end to end. The demo arena has five pre-forged bosses, no upload needed.

photo → 3D boss · arena combat · 2–4 min per forge

bossforge-one.vercel.app

05

Chitra

Take a selfie, get every photo of yourself out of a ten-thousand-image event gallery.

Event galleries are where photos go to disappear: thousands of images, no names, no order. Chitra searches them by face. One selfie in, every photo you appear in out. I built it alone, idea to production: the matching pipeline, the infrastructure, the product.

face matching · solo, idea → production

chitra.cloud

t≈400 · publications

LSTMSE-Net: Long Short Term Speech Enhancement Network for Audio-visual Speech Enhancement

Fuses visual cues with audio through an encoder-decoder separator network to clean up speech.

INTERSPEECH 2024arXiv:2409.02266

AUREXA-SE: Audio-Visual Unified Representation Exchange Architecture with Cross-Attention and Squeezeformer for Speech Enhancement

Exchanges audio and visual representations through cross-attention, with Squeezeformer blocks doing the temporal modeling.

AVSEC Workshop, INTERSPEECH 2025arXiv:2510.05295

t≈150 · ledger

Silver
Inter IIT Tech Meet 14.0, 2025
Silver
IITISoC, IIT Indore, 2024
3rd
Databricks Hackathon (IIT-level), 2025
1st
Enosium Hackathon, Fluxus, IIT Indore, 2024
200+
problems on Codeforces and LeetCode
President
Cynaptics Club, IIT Indore, 2025–26
Coordinator
UG SCC, Counseling Cell, IIT Indore, 2024–25
8.27
CGPA, B.Tech Electrical Engineering, IIT Indore

t=0 · signal

Building something that has to be fast?

harshvardhan.yashvardhan@gmail.com

GitHubLinkedInRésumé