vLLM
High-throughput open-source LLM inference and serving engine.
Pricing
Open Source
Category
Local AI
Open source
Yes
In the database since
2026-07-01
Overview
About vLLM
High-throughput open-source LLM inference and serving engine.
Best for: Production-grade self-hosted model serving and throughput.
Not ideal for: Casual desktop chat without GPU ops experience.
Best for
Production-grade self-hosted model serving and throughput.
Not ideal for
Casual desktop chat without GPU ops experience.
Keep exploring
Similar AI tools
Related by category and capabilities.
Run open models locally with a simple CLI and OpenAI-compatible API.
Local GGUF inference with creative writing and roleplay-oriented features.
Self-hosted chat UI for Ollama and OpenAI-compatible backends.
Desktop app for downloading and chatting with local GGUF models.
Workflows using vLLM · 9 verified
View all →Daily Slack Channel Summarizer With n8n And Claude
Automatically summarize key decisions, action items, and open questions from Slack channels into a daily digest using Claude 3.5 Sonnet.
AI Social Media Content Scheduler With n8n And Open AI
Auto-generate and schedule LinkedIn, Twitter, and Instagram posts from your Notion content calendar using GPT-4o and Buffer.
AI Customer Support Ticket Escalation With n8n And Open AI
Classify Zendesk tickets by urgency with GPT-4o and auto-route: auto-reply low, queue medium, Slack+PagerDuty for critical.
AI Email Inbox Auto Responder With Make.com And Chat GPT
Download the Make.com blueprint to build an automated email answering system that drafts context-aware replies using ChatGPT.
Turn Any Podcast Or Meeting Transcript Into A Blog Post With Make.com
Convert any podcast transcript or meeting recording into a full SEO blog post automatically using GPT-4o and Make.com.
AI Cold Email Personalizer With n8n And GPT 4o
Automatically write hyper-personalized cold emails for each prospect using GPT-4o, pulling live company data and recent news via n8n.
Automated Competitor Price And Feature Monitor With Make.com
Scrape competitor pricing pages weekly with Apify, detect changes using GPT-4o, and alert your team in Slack before they alert your customers.
Multi Platform Content Repurposer With n8n And Claude
Turn a single blog post into LinkedIn posts, Twitter threads, and newsletter intros automatically using Claude and n8n.
AI Contract Risk Analyzer With Make.com And Claude
Analyze PDF contracts with Claude to identify risk clauses, missing terms, and payment conditions � with plain-English summaries.
Related MCP & skills
Web content fetching and conversion for efficient LLM usage.
Node.js server implementing Model Context Protocol (MCP) for filesystem operations.
This repository is a collection of reference implementations for the [Model Context Protocol](https://modelcontextprotocol.io/) (MCP), as well as references…
A basic implementation of persistent memory using a local knowledge graph. This lets Claude remember information about the user across chats.
Dynamic and reflective problem-solving through thought sequences.
📇 ☁️ - A MCP server for the Open Library API that enables AI assistants to search for book information.