【文章标题】:WebLLM: 高性能浏览器内大语言模型推理引擎

【文章正文】: High-Performance In-Browser LLM Inference Engine. 高性能浏览器内大语言模型推理引擎 Documentation | Blogpost | Paper | Examples 文档 | 博客文章 | 论文 | 示例

WebLLM is a high-performance in-browser LLM inference engine that brings language model inference directly onto web browsers with hardware acceleration. Everything runs inside the browser with no server support and is accelerated with WebGPU. WebLLM是一个高性能的浏览器内大语言模型推理引擎,通过硬件加速将语言模型推理直接引入网页浏览器。所有操作均在浏览器内运行,无需服务器支持,并通过WebGPU实现加速。

WebLLM is fully compatible with OpenAI API. That is, you can use the same OpenAI API on any open source models locally, with functionalities including streaming, JSON-mode, function-calling (WIP), etc. WebLLM完全兼容OpenAI API。这意味着您可以在本地任何开源模型上使用相同的OpenAI API,功能包括流式传输、JSON模式、函数调用(开发中)等。

We can bring a lot of fun opportunities to build AI assistants for everyone and enable privacy while enjoying GPU acceleration. 我们能为每个人创造构建AI助手的趣味机会,在享受GPU加速的同时保障隐私安全。

You can use WebLLM as a base npm package and build your own web application on top of it by following the examples below. This project is a companion project of MLC LLM, which enables universal deployment of LLM across hardware environments. 您可以将WebLLM作为基础npm包使用,并参照以下示例构建自己的网页应用。该项目是MLC LLM的配套项目,实现大语言模型在各类硬件环境中的通用部署。

In-Browser Inference: WebLLM is a high-performance, in-browser language model inference engine that leverages WebGPU for hardware acceleration, enabling powerful LLM operations directly within web browsers without server-side processing.

  • 浏览器内推理:WebLLM作为高性能的浏览器内语言模型推理引擎,利用WebGPU进行硬件加速,无需服务器端处理即可在浏览器内直接执行强大的LLM运算。

Full OpenAI API Compatibility: Seamlessly integrate your app with WebLLM using OpenAI API with functionalities such as streaming, JSON-mode, logit-level control, seeding, and more.

  • 完全OpenAI API兼容:通过支持流式传输、JSON模式、逻辑控制、种子设置等功能的OpenAI API,实现应用与WebLLM的无缝集成。

Structured JSON Generation: WebLLM supports state-of-the-art JSON mode structured generation, implemented in the WebAssembly portion of the model library for optimal performance. Check WebLLM JSON Playground on HuggingFace to try generating JSON output with custom JSON schema.

  • 结构化JSON生成:WebLLM支持最先进的JSON模式结构化生成,该功能在模型库的WebAssembly部分实现以获得最佳性能。访问HuggingFace上的WebLLM JSON Playground,尝试使用自定义JSON模式生成JSON输出。

Extensive Model Support: WebLLM natively supports a range of models including Llama 3, Phi 3, Gemma, Mistral, Qwen(通义千问), and many others, making it versatile for various AI tasks. For the complete supported model list, check MLC Models.

  • 广泛模型支持:WebLLM原生支持包括Llama 3、Phi 3、Gemma、Mistral、Qwen(通义千问)等在内的多种模型,适用于各类AI任务。完整支持模型列表请查看MLC Models。

Custom Model Integration: Easily integrate and deploy custom models in MLC format, allowing you to adapt WebLLM to specific needs and scenarios, enhancing flexibility in model deployment.

  • 自定义模型集成:轻松集成并部署MLC格式的自定义模型,使WebLLM能适应特定需求和场景,提升模型部署的灵活性。

Plug-and-Play Integration: Easily integrate WebLLM into your projects using package managers like NPM and Yarn, or directly via CDN, complete with comprehensive examples and a modular design for connecting with UI components.

  • 即插即用集成:通过NPM/Yarn等包管理器或直接通过CDN轻松将WebLLM集成至项目,提供完整示例和模块化设计以便连接UI组件。

Streaming & Real-Time Interactions: Supports streaming chat completions, allowing real-time output generation which enhances interactive applications like chatbots and virtual assistants.

  • 流式与实时交互:支持流式聊天补全功能,实现实时输出生成,提升聊天机器人和虚拟助手等交互式应用的体验。

Web Worker & Service Worker Support: Optimize UI performance and manage the lifecycle of models efficiently by offloading computations to separate worker threads or service workers.

  • Web Worker与服务Worker支持:通过将计算任务卸载至独立worker线程或service worker,优化UI性能并高效管理模型生命周期。

Chrome Extension Support: Extend the functionality of web browsers through custom Chrome extensions using WebLLM, with examples available for building both basic and advanced extensions.

  • Chrome扩展支持:使用WebLLM构建自定义Chrome扩展来增强浏览器功能,提供基础版和高级版扩展的构建示例。

Check the complete list of available models on MLC Models. WebLLM supports a subset of these available models and the list can be accessed at prebuiltAppConfig.model_list. 完整模型列表请查看MLC Models。WebLLM支持其中部分模型,具体列表可通过prebuiltAppConfig.model_list获取。

Here are the primary families of models currently supported: 当前主要支持的模型系列包括:

  • Llama: Llama 3, Llama 2, Hermes-2-Pro-Llama-3
  • Phi: Phi 3, Phi 2, Phi 1.5
  • Gemma: Gemma-2B
  • Mistral: Mistral-7B-v0.3, Hermes-2-Pro-Mistral-7B, NeuralHermes-2.5-Mistral-7B, OpenHermes-2.5-Mistral-7B
  • Qwen (通义千问): Qwen2 0.5B, 1.5B, 7B

If you need more models, request a new model via opening an issue or check Custom Models for how to compile and use your own models with WebLLM. 如需更多模型,可通过提交issue申请新模型,或参考Custom Models了解如何编译和使用自定义模型。

Learn how to use WebLLM to integrate large language models into your application and generate chat completions through this simple Chatbot example: 通过这个简单的聊天机器人示例,了解如何使用WebLLM将大语言模型集成到应用中并生成聊天补全:

For an advanced example of a larger, more complicated project, check WebLLM Chat. 更复杂项目的进阶示例请查看WebLLM Chat。

More examples for different use cases are available in the examples folder. 更多用例示例可在examples文件夹中获取。

WebLLM offers a minimalist and modular interface to access the chatbot in the browser. The package is designed in a modular way to hook to any of the UI components. WebLLM提供极简模块化接口来访问浏览器中的聊天机器人。该包采用模块化设计,可连接任何UI组件。

npm

npm install @mlc-ai/web-llm

yarn

yarn add @mlc-ai/web-llm

or pnpm

pnpm install @mlc-ai/web-llm

Then import the module in your code. 然后在代码中导入模块: // Import everything import * as webllm from “@mlc-ai/web-llm”; // Or only import what you need import { CreateMLCEngine } from “@mlc-ai/web-llm”;

Thanks to jsdelivr.com, WebLLM can be imported directly through URL and work out-of-the-box on cloud development platforms like jsfiddle.net, Codepen.io, and Scribbler: 感谢jsdelivr.com,WebLLM可直接通过URL导入,并在jsfiddle.net、Codepen.io和Scribbler等云开发平台上开箱即用: import * as webllm from “https://esm.run/@mlc-ai/web-llm”;

It can also be dynamically imported as: 也可动态导入: const webllm = await import(“https://esm.run/@mlc-ai/web-llm”);

Most operations in WebLLM are invoked through the MLCEngine interface. You can create an MLCEngine instance and loading the model by calling the CreateMLCEngine() factory function. WebLLM中的大多数操作通过MLCEngine接口调用。您可以通过CreateMLCEngine()工厂函数创建MLCEngine实例并加载模型。

(Note that loading models requires downloading and it can take a significant amount of time for the very first run without caching previously. You should properly handle this asynchronous call.) (注意:首次无缓存加载模型需要下载,可能耗费较长时间。您应妥善处理此异步调用。)

🔗 知识库双向关联