My Journey into Ethical Web Scraping with Python
Best practices for responsible web scraping, including respecting robots.txt, rate limiting, and legal considerations when extracting data from websites.
I'm Abdur Rafay, an AI Engineer at Lythe Labs. I build production-grade LLM applications, RAG pipelines, and multi-agent systems, with experience shipping 25+ AI agents across healthcare, finance, and education.
Get to know the person behind the code
I'm an AI Engineer currently working at Lythe Labs (Singapore, Remote) building production-grade LLM applications, RAG pipelines, and multi-agent systems. I've shipped 25+ AI agents across healthcare, finance, and education domains.
I'm pursuing my BS in Artificial Intelligence at Bahria University while working across 5+ international internships, from Cairo to Georgia, Dubai to Lahore. I thrive at the intersection of research and engineering: taking ideas from paper to production.
BS Artificial Intelligence, Bahria University (2024-Present)
Wah Cantt, Pakistan
LLMs, Multi-Agent Systems, RAG Pipelines
What I can do for you
Building intelligent systems using machine learning and deep learning. Experienced with LLMs, prompt engineering, fine-tuning, and working with TensorFlow and PyTorch.
Creating modern web applications using Flask, Django, React, and modern frontend technologies. From backend APIs to responsive user interfaces.
Expert in web scraping and automation using Python libraries like Requests, BeautifulSoup, and Selenium for data extraction and process automation.
Technologies I work with
Primary language for AI/ML, backend APIs, automation, and data pipelines
Complex queries, schema design, and optimization across PostgreSQL and MySQL
Frontend scripting, async patterns, and Node.js tooling
Building agentic pipelines, RAG systems, and tool-augmented LLM applications
Sentence Transformers, embeddings, and open-source model fine-tuning
Deep learning, CNNs, LSTMs, and classical ML with Scikit-learn/XGBoost
OpenAI, Anthropic Claude, Groq, Gemini: prompt engineering and fine-tuning
Production RAG pipelines, vector search, and orchestrated multi-agent workflows
High-performance async REST APIs for AI services and data pipelines
Lightweight APIs and rapid AI/data app prototyping and deployment
Modern frontend development with fast builds and component-driven UIs
Relational databases: schema design, indexing, and analytical queries
Vector database for embedding storage and semantic similarity search in RAG systems
Containerised deployments, CI/CD pipelines, and version control workflows
My journey so far
Building production-grade AI agents and multi-agent systems at a stealth-stage AI startup. Architecting LLM pipelines, RAG systems, and autonomous agent workflows deployed across multiple domains.
Developed and fine-tuned LLM-based applications, contributed to curriculum content on prompt engineering and agent design, and collaborated with a global team of AI researchers and engineers.
Built RAG pipelines and LLM-powered automation tools for internal business workflows. Worked with ChromaDB, LangChain, and FastAPI to deliver measurable efficiency improvements.
Applied machine learning models to learning pathway personalisation. Implemented Scikit-learn and XGBoost pipelines for recommendation and classification tasks on educational data.
Performed data analysis and reporting using SQL and Python. Built dashboards and automated data pipelines to surface actionable insights for business stakeholders.
Developed ML models for cybersecurity data analysis. Built anomaly detection pipelines using Python, Pandas, and Scikit-learn to identify patterns in security event logs.
Studying AI fundamentals through to advanced deep learning, NLP, computer vision, and intelligent systems design alongside real-world engineering work.
Some of my recent work
Self-hosted AI agent capable of running autonomously for days or weeks. Uses LLM-driven planning loops, sandboxed Python execution, persistent memory, and resumable sessions via API control. Structured planning minimises hallucination rates.
Thoughts and insights on AI, development, and technology
Best practices for responsible web scraping, including respecting robots.txt, rate limiting, and legal considerations when extracting data from websites.
An overview of practical lessons learned while implementing Retrieval-Augmented Generation (RAG) systems in a startup environment, highlighting how retrieval quality, data preprocessing, and system design impact the performance and reliability of LLM-based applications.
Professional certifications and course completions
This comprehensive professional certificate program covers the complete journey of AI-powered software development. It includes foundational knowledge in software engineering, artificial intelligence, and generative AI, progressing through practical application development using Python, Flask, HTML, CSS, and JavaScript.
This NVIDIA course provides foundational knowledge in AI infrastructure and operations management. It covers the essential concepts, tools, and best practices for deploying, managing, and operating AI systems at scale, focusing on the technical infrastructure required to support AI workloads effectively.
I take on freelance projects through Upwork: LLM apps, RAG pipelines, multi-agent systems, and AI integrations. Let's build something together.
Let's connect and discuss opportunities
I'm always interested in hearing about new projects and opportunities. Whether you have a question or just want to say hi, feel free to reach out!