Browsing properties or booking a stay on Booking.com triggers a massive real-time machinery of over 480 machine learning models running simultaneously behind the scenes. Operating at a scale that processes anywhere from 500 billion to 800 billion predictions per day with 99.99% reliability, maintaining sub-20-millisecond latency while integrating non-deterministic Large Language Models (LLMs) into production presents a unique engineering challenge.
At the 2025 edition of the Nordic Data Science and Machine Learning (NDSML) Summit, Sanchit Juneja, Product Lead for Big Data and Machine Learning at Booking.com, shared more about “LLMOps: Factory Floor for GenAI in Booking.com”. His session detailed how the travel giant evolved an industry-leading MLOps ecosystem into a specialized LLMOps architecture designed to handle generative AI at enterprise scale.
ROI Over Research
Before diving into hardware and pipelines, Juneja established a clear distinction between academic AI research and applied enterprise engineering. While research institutions focus on foundational model training, for-profit organizations focus on operational execution: FM Ops (Foundation Model Operations) and its subset, LLM Ops.
Just as software engineering created DevOps and security yielded SecOps, the complexity of managing live generative models requires a dedicated tooling layer: the “factory floor” that turns AI research into measurable business revenue.
From MLOps to LLMOps
In traditional machine learning, production systems rely heavily on structured data, self-hosted models, and deterministic outputs. As Juneja explained during the presentation, the rise of LLMs forced a fundamental architectural shift across enterprise environments:
- Unstructured Data Ingestion
Production stacks must now seamlessly process direct user text, audio, and image uploads rather than relying strictly on formatted protocol buffers.
- LLM-as-a-Service & Middleware Wrappers
Not every business problem requires self-hosted models. Third-party hosted models can be leveraged through customized middleware layers. They can provide domain-agnostic access across external vendors without compromising compliance or performance.
- Deterministic Guardrails
Serving travelers urgent updates requires speed and precision. Implementing specialized Prompt Stores, Vector Stores, and Graph Neural Networks enforces determinism and low latency over non-deterministic base models.
- Hardware Evolution
Scaled GenAI requires distributed GPU inference alongside specialized hardware accelerators like XPUs and Groq to maintain real-time performance.
Governance, Feedback Loops, and the Rise of SLMs
Deploying models at global scale requires rigorous oversight. Compliance with evolving governance frameworks (such as the EU AI Act) demands continuous monitoring for model drift, toxicity, and algorithmic bias. To maintain quality, Booking.com leverages automated feedback loops. It incorporates Reinforcement Learning from Human Feedback (RLHF) to continuously retrain stale models and streamline label collection.
Looking toward the future, Juneja highlighted a shift away from ever-larger foundational models toward Small Language Models (SLMs). Hyper-specialized 1B to 3B parameter models offer light compute footprints, making them ideal for targeted enterprise tasks and edge deployments.
Go Deep Into the Production Architecture
From evaluating models using “LLM-as-a-Judge” frameworks to fine-tuning targeted Small Language Models (SLMs) and managing RLHF feedback loops, the presentation provides the operational plan needed to move from generative hype to high-ROI production systems.
This overview barely scratches the surface of the full operational playbook. In the complete session recording, Sanchit Juneja breaks down the exact schematic of Booking.com’s model serving layer, data lake orchestration, guardrailing strategies, and GPU cost-optimization tools like Run:ai.
Access the full video to explore the exact technical details and see how top-tier engineering teams construct production-ready LLM ecosystems.
Join the Next NDSML Summit
With capacity strictly capped at 400 delegates, and over 70% of onsite places already booked, enterprise teams are now finalizing their plans for this year’s edition of the NDSML Event.
Discover cutting-edge strategies directly from world-class data scientists, MLOps architects, and AI pioneers in person. Reserve a spot for the upcoming NDSML Summit in Stockholm to participate in technical tracks, executive panels, and networking sessions with leaders driving the future of enterprise AI.