Languages
- PythonAgents, backend APIs, automation, data pipelines
- SQLComplex queries, schema design, optimisation
- JavaScriptFrontend scripting, async patterns, Node tooling
- C++Systems fundamentals
Available for freelance & full-time roles
I'm Abdur Rafay, Member of Technical Staff at Lythe Labs (San Francisco, remote). I ship production voice agents against a sub‑800 ms time-to-first-audio budget: streaming ASR → LLM → TTS, client-side VAD, barge-in, fully async. I own the CI/CD and AWS stack underneath them too. 25+ agent systems delivered across healthcare, finance, cybersecurity & education.
Most AI work fails somewhere other than the prompt: in retrieval quality, in a blocking call that adds half a second to a phone conversation, in the deploy that nobody automated. I work on that whole surface, not just the model call.
At Lythe Labs I build production voice agents: real-time pipelines that stream ASR → LLM → TTS over a WebSocket so audio starts playing before the sentence is finished, client-side VAD that decides when a caller has actually stopped talking and cuts playback the moment they start speaking over it, and the tool-calling agent infrastructure underneath. The open-source Voice Agent Starter Kit came out of that work. I also own the parts nobody volunteers for: CI/CD with automated test integration, the AWS production stack, MCP server development, and test plans for telephony-side tooling.
Before this I delivered 25+ agent systems end-to-end across healthcare, finance, cybersecurity and education, and built RAG pipelines in production. I'm completing a BS in Artificial Intelligence at Bahria University alongside the work, and I publish the tools I need but can't find: an MCP latency auditor, a dependency-free agent framework, an LLM cache.
Streaming ASR → LLM → TTS pipelines, WebSocket audio transport, VAD, endpointing and barge-in, latency budgets and telemetry.
Tool-calling and multi-agent orchestration, MCP servers, memory and retrieval layers, evaluation harnesses.
CI/CD pipelines with automated tests, AWS and GCP production stacks, Docker, release and deploy ownership.
FastAPI and Flask services, PostgreSQL, MCP servers, and the API surface agents actually call.
Test plans, pytest and API testing, regression suites in CI, defect triage through Jira and Linear.
If it is on this list, it has shipped in something real.
Every entry links to real source. One of them you can pip install.
A complete real-time voice loop, microphone to spoken reply, with the latency work already done and the ASR, LLM and TTS providers left swappable.
main.py.An MCP server that audits Python voice-agent codebases for the async mistakes that quietly add latency to a phone call.
relay-arclat. Latency review becomes an automated pass instead of a careful read-through.A multi-agent framework that runs entirely on your machine: hybrid retrieval over SQLite, local models through Ollama, zero core dependencies.
An evaluation dashboard that makes multi-agent pipelines legible: where the time goes, where the money goes, which tools actually get called.
Automatic caching of LLM responses across OpenAI, Anthropic and Gemini behind one interface. Repeat calls resolve locally, cutting spend and latency on repeated NLP queries.
Autonomous agent built for weeks of continuous runtime: LLM planning loops, sandboxed code execution, persistent memory and resumable sessions that survive a restart.
Security scanner for MCP servers, in the shape of npm audit. Scans the servers you install and the ones you build for tool poisoning, over-broad permissions, credential exposure and rug pulls. Zero runtime dependencies.
Every role has come down to the same thing: make the model useful, then keep it running.
Building and shipping production voice agents, multi-agent systems, and the infrastructure that runs them: system design through deployment, solo, across a distributed remote team.
Delivered 25+ AI agent systems independently across healthcare, finance, cybersecurity and education, each end-to-end with minimal supervision.
Built and optimised production RAG pipelines on Groq and ChromaDB, holding low latency at scale without losing retrieval accuracy.
Six distinct ML projects delivered independently in one month, each through the full lifecycle: cleaning, feature engineering, training, evaluation and visualisation.
Led a team of four interns analysing social media marketing data.
Data work and QA for a security analytics team.
Bahria University, Islamabad · 2024 - Present
AI fundamentals through to deep learning, NLP, computer vision and intelligent systems design, studied alongside full-time engineering work.
Transformer architecture, model selection, fine-tuning and RLHF, and the deployment considerations behind LLM-powered applications.
Concepts, tooling and practice for deploying, managing and operating AI systems at scale, covering the infrastructure side of supporting AI workloads.
VerifyFoundations of supervised learning: linear and logistic regression, gradient descent, regularisation and model evaluation.
The full journey of AI-powered software development: software engineering and generative AI foundations through to shipping applications with Python and Flask.
Verify
AI & machine learning
Practical lessons from implementing retrieval-augmented generation in a startup: how retrieval quality, preprocessing and system design decide whether an LLM app is reliable.
Read on Medium
Web scraping
Respecting robots.txt, rate limiting and the legal considerations that separate responsible data collection from the kind that gets you blocked.
Read on MediumOpen to full-time AI engineering roles and freelance work: real-time voice agents, agent infrastructure, RAG pipelines, and the cloud stack to run them. Tell me what you're building.
abdurrafay.tech@gmail.comCurrently available · replies within 24 hours