Islam Kassem

AI Engineer · LLM & Generative AI · Kaggle Code Master (Top 1%) · MSc in Data Science

Building production AI systems that deliver measurable business value.

I design and deploy generative AI systems, LLM pipelines, RAG search, Whisper ASR, and multimodal vision — with a strong focus on inference efficiency, scalable MLOps, and production-ready delivery on AWS and Kubernetes.

About

AI Engineer shipping LLM, speech, and vision systems into production.

I work across the full AI delivery stack: data pipelines, experimentation, fine-tuning, inference optimization, deployment, and monitoring. Today that means enterprise RAG, on-premise LLMs, and speech-to-text pipelines for organizations that need their data to stay in-house.

I care about systems that are reliable, cost-aware, and tied to business outcomes rather than demos: LoRA/QLoRA fine-tuning of open-weight LLMs, retrieval over FAISS and vector databases, real-time inference, and MLOps with CI/CD, MLflow, and experiment tracking.

Skills

Core capabilities across model development, systems, and cloud delivery.

AI / ML

LLMs RAG AI Agents LoRA / QLoRA Whisper ASR Computer Vision Multimodal AI Semantic Search NLP Recommendation Systems PyTorch Hugging Face LangChain

Serving / Backend

Python FastAPI vLLM Ollama REST APIs gRPC Kafka SQL MongoDB FAISS Pinecone Elasticsearch

Cloud / MLOps

AWS GCP Vertex AI Azure Docker Kubernetes GitHub Actions MLflow Weights & Biases Model Monitoring

Projects

Selected work across enterprise AI and live ML products.

On-Premise LLM & Speech Platform

Enterprise RAG assistants and Whisper large-v3 transcription running on self-hosted LLaMA and Qwen models on GPU Kubernetes, so sensitive data never leaves the client's infrastructure.

Impact: data-sovereign generative AI for privacy-critical environments.

On-Prem LLMs RAG Whisper ASR Kubernetes

AI Meeting Assistant

A multi-stage pipeline that transcribes meetings with Whisper, summarizes them with LLMs, and uses LangChain agents to extract action items and draft follow-ups.

Impact: post-meeting documentation cut from hours to minutes.

Whisper LLM Agents Summarization

Semantic Resume–Role Matching

Dense retrieval with sentence-transformers and FAISS, followed by re-ranking, to match candidate resumes to job requirements across thousands of applications.

Impact: deployed in government hiring workflows.

Embeddings FAISS Re-ranking

OpenFPL Scout AI

OpenFPL Scout AI is a live Fantasy Premier League scout that pairs historical form with official FPL data to forecast every player, optimize a 15-player squad, and rank captaincy picks before each gameweek deadline.

Impact: a free public app, rebuilt every gameweek from live data.

Ensemble Models Time Series Optimization Live Data

Services

Productized AI and ML services available on Upwork.

AI / ML Consultation

One-on-one technical consultation to scope AI opportunities, architecture decisions, and practical implementation plans.

Consultation Solution Design Roadmapping

Custom LLM Agents for Business

Design and development of custom LLM-powered agents tailored to automate workflows and support domain-specific tasks.

LLMs AI Agents Automation

Credentials

Education, certifications, and recognitions.

Education

M.Sc. in Data Science, Cairo University (in progress, since 2024) · B.Sc. in Electrical and Communication Engineering, Port Said University (2021)

Certifications

Deep Learning Specialization (DeepLearning.AI) · Analyze Datasets and Train ML Models using AutoML (AWS)

Recognitions

Kaggle Code Master (Top 1%), 2 silver and 2 bronze medals · Upwork Top-Rated, 100% Job Success Score

Blog

Latest articles.

OpenFPL Scout AI: What a Full Season of Blind Validation Actually Proved

A strict out-of-sample test of OpenFPL: train on two seasons, score once on a full season the models never saw, and report honestly what the forecasts can and can't do.

Aug 2026 · Machine Learning · Model Evaluation

Machine Learning Model Evaluation Fantasy Premier League

Understanding Large Language Models: A Beginner's Guide

A ground-up explanation of how Large Language Models work, covering transformer architecture, pretraining, fine-tuning, and the intuition behind coherent text generation.

Feb 2026 · Large Language Models · Generative AI

LLMs Generative AI Deep Learning

Contact

Open to consulting, collaborations, and select AI engineering opportunities.

If you need LLM systems, semantic retrieval, speech pipelines, or ML platform work shipped with production discipline, reach out directly.