Media Summary: In this session, we explored the motivation for Timestamps: 00:00 - Intro 01:24 - Technical Demo 09:48 - Results 11:02 - Intermission 11:57 - Considerations 15:48 - Conclusion ... Running Large Language Models (LLMs) locally for experimentation is easy but running them in large scale architectures is not.

Vllm Office Hours Distributed Inference - Detailed Analysis & Overview

In this session, we explored the motivation for Timestamps: 00:00 - Intro 01:24 - Technical Demo 09:48 - Results 11:02 - Intermission 11:57 - Considerations 15:48 - Conclusion ... Running Large Language Models (LLMs) locally for experimentation is easy but running them in large scale architectures is not. In this session, we explored the latest updates in the This walkthrough showcases how to deploy large language model (LLM) Ready to become a certified watsonx AI Assistant Engineer? Register now and use code IBMTechYT20 for 20% off of your exam ...

Photo Gallery

vLLM Office Hours - Distributed Inference with vLLM - January 23, 2025
[vLLM Office Hours #48] vLLM Project and Tool Calling Update - April 30, 2026
Run A Local LLM Across Multiple Computers! (vLLM Distributed Inference)
[vLLM Office Hours #49] Latest Trends in AI Agent Applications and vLLM - May 18, 2026
[vLLM Office Hours #53] - llm-d Project Update and Wide EP for Agentic Workloads - July 9, 2026
[vLLM Office Hours #55] - Mooncake + vLLM/llm-d Deep Dive - August 6, 2026
Large Scale Distributed LLM Inference with LLM D and Kubernetes by Abdel Sghiouar
[vLLM Office Hours #39] Intro to batch invariant in vLLM - January 8, 2026
[vLLM Office Hours #45] Vienna vLLM Meetup Live Stream - March 11, 2026
[vLLM Office Hours #27] Intro to llm-d for Distributed LLM Inference
Distributed LLM inferencing across virtual machines using vLLM and Ray
What is vLLM? Efficient AI Inference for Large Language Models
View Detailed Profile
vLLM Office Hours - Distributed Inference with vLLM - January 23, 2025

vLLM Office Hours - Distributed Inference with vLLM - January 23, 2025

In this session, we explored the motivation for

[vLLM Office Hours #48] vLLM Project and Tool Calling Update - April 30, 2026

[vLLM Office Hours #48] vLLM Project and Tool Calling Update - April 30, 2026

... Tool Calling in

Run A Local LLM Across Multiple Computers! (vLLM Distributed Inference)

Run A Local LLM Across Multiple Computers! (vLLM Distributed Inference)

Timestamps: 00:00 - Intro 01:24 - Technical Demo 09:48 - Results 11:02 - Intermission 11:57 - Considerations 15:48 - Conclusion ...

[vLLM Office Hours #49] Latest Trends in AI Agent Applications and vLLM - May 18, 2026

[vLLM Office Hours #49] Latest Trends in AI Agent Applications and vLLM - May 18, 2026

Welcome to

[vLLM Office Hours #53] - llm-d Project Update and Wide EP for Agentic Workloads - July 9, 2026

[vLLM Office Hours #53] - llm-d Project Update and Wide EP for Agentic Workloads - July 9, 2026

Welcome to

[vLLM Office Hours #55] - Mooncake + vLLM/llm-d Deep Dive - August 6, 2026

[vLLM Office Hours #55] - Mooncake + vLLM/llm-d Deep Dive - August 6, 2026

Welcome to

Large Scale Distributed LLM Inference with LLM D and Kubernetes by Abdel Sghiouar

Large Scale Distributed LLM Inference with LLM D and Kubernetes by Abdel Sghiouar

Running Large Language Models (LLMs) locally for experimentation is easy but running them in large scale architectures is not.

[vLLM Office Hours #39] Intro to batch invariant in vLLM - January 8, 2026

[vLLM Office Hours #39] Intro to batch invariant in vLLM - January 8, 2026

In this

[vLLM Office Hours #45] Vienna vLLM Meetup Live Stream - March 11, 2026

[vLLM Office Hours #45] Vienna vLLM Meetup Live Stream - March 11, 2026

Tune in to the Vienna

[vLLM Office Hours #27] Intro to llm-d for Distributed LLM Inference

[vLLM Office Hours #27] Intro to llm-d for Distributed LLM Inference

In this session, we explored the latest updates in the

Distributed LLM inferencing across virtual machines using vLLM and Ray

Distributed LLM inferencing across virtual machines using vLLM and Ray

This walkthrough showcases how to deploy large language model (LLM)

What is vLLM? Efficient AI Inference for Large Language Models

What is vLLM? Efficient AI Inference for Large Language Models

Ready to become a certified watsonx AI Assistant Engineer? Register now and use code IBMTechYT20 for 20% off of your exam ...

[vLLM Office Hours #28] GuideLLM: Evaluate your LLM Deployments for Real-World Inference

[vLLM Office Hours #28] GuideLLM: Evaluate your LLM Deployments for Real-World Inference

In this bi-weekly session of the