用 LlamaIndex 和 llamafile 搭建本地私有研究助手
Using LlamaIndex and llamafile to build a local, private research assistant
LlamaIndex 发布 Mozilla 客座教程,讲解用 llamafile 在笔记本上本地运行 LLM,并以 TinyLlama-1.1B 作为 LLM 与嵌入后端,结合 Wikipedia 页面与私有文档构建完全本地运行的 RAG 研究助手。
教程给出从下载运行 llamafile 到用 LlamaIndex 搭建本地 RAG 研究助手的完整代码路径,可迁移到其他主题。
这是一篇来自我们的朋友 Mozilla 的客座文章,内容关于 Llamafile
llamafile 是 Mozilla 的一个开源项目,是在笔记本电脑上运行大型语言模型(LLM)最简单的方法之一。你只需从 HuggingFace 下载一个 llamafile,然后运行该文件即可。就这么简单。在大多数电脑上,你无需安装任何东西。
你可能想在笔记本电脑上运行 LLM 的原因有几点,包括:
1. 隐私:本地运行意味着你无需与第三方共享数据。
2. 高可用性:无需互联网连接即可运行基于 LLM 的应用。
3. 自带模型:你可以轻松测试许多不同的开源 LLM(HuggingFace 上可用的任何模型),看看哪一个最适合你的任务。
4. 免费调试/测试:本地 LLM 允许你测试基于 LLM 系统的许多部分,而无需为 API 调用付费。
在这篇博客文章中,我们将展示如何设置 llamafile 并使用它在你的电脑上运行本地 LLM。然后,我们将展示如何将 LlamaIndex 与你的 llamafile 结合使用,作为本地基于 RAG 的研究助手的 LLM 与嵌入后端。你无需注册任何云服务,也无需将数据发送给任何第三方——一切都将在你的笔记本电脑上运行。
注意:你也可以从我们的 GitHub 仓库 以 Jupyter notebook 的形式获取下面所有的示例代码。
立即探索我们的免费和付费方案。
下载并运行 llamafile
首先,什么是 llamafile?llamafile 是一个可执行文件形式的 LLM,你可以在自己的电脑上运行。它包含给定开源 LLM 的权重,以及在你的电脑上实际运行该模型所需的一切。无需安装或配置任何东西(有一些注意事项,详见此处)。
每个 llamafile 都打包了 1) gguf 格式的模型权重与元数据 + 2) 一份使用 [Cosmopolitan Libc](https://github.com/jart/cosmopolitan) 特别编译的 `llama.cpp` 副本。这使得这些模型可以在大多数电脑上运行,无需额外安装。llamafiles 还附带类似 ChatGPT 的浏览器界面、CLI,以及用于聊天模型的 OpenAI 兼容 REST API。
设置 llamafile 只需 2 个步骤:
1. 下载 llamafile
2. 使 llamafile 可执行
我们将在下面详细介绍每个步骤。
步骤 1:下载 llamafile
HuggingFace 模型中心 上有许多可用的 llamafiles(只需搜索 'llamafile'),但为了本演练的目的,我们将使用 TinyLlama-1.1B(0.67 GB,模型信息)。要下载该模型,你可以点击此下载链接:TinyLlama-1.1B,或者打开终端并使用类似 `wget` 的命令。下载大约需要 5-10 分钟,具体取决于你的互联网连接质量。
wget https://huggingface.co/Mozilla/TinyLlama-1.1B-Chat-v1.0-llamafile/resolve/main/TinyLlama-1.1B-Chat-v1.0.F16.llamafile 这个模型很小,实际上并不擅长回答问题,但由于它的下载相对较快,而且其推理速度可以让你在几分钟内完成向量存储的索引,因此对于下面的示例来说已经足够好了。如果你想要更高质量的 LLM,可能需要使用更大的模型,例如 Mistral-7B-Instruct(5.15 GB,模型信息)。
步骤 2:使 llamafile 可执行
如果你不是从命令行下载的 llamafile,请找出你的浏览器将下载的 llamafile 存储在了哪里。
现在,打开你的电脑终端,如有必要,进入你存储 llamafile 的目录:`cd path/to/downloaded/llamafile`
如果你使用的是 macOS、Linux 或 BSD,你需要授予你的电脑执行这个新文件的权限。(你只需要做一次):
如果你使用的是 Windows,只需在文件名末尾加上 ".exe" 来重命名文件,例如将 `TinyLlama-1.1B-Chat-v1.0.Q5_K_M.llamafile` 重命名为 `TinyLlama-1.1B-Chat-v1.0.Q5_K_M.llamafile.exe`
chmod +x TinyLlama-1.1B-Chat-v1.0.Q5_K_M.llamafile试驾一下
现在,你的 llamafile 应该已经准备就绪了。首先,你可以检查一下你下载的 llamafile 二进制文件是用哪个版本的 llamafile 库构建的:
./TinyLlama-1.1B-Chat-v1.0.Q5_K_M.llamafile --version
llamafile v0.7.0这篇文章是使用 `llamafile v0.7.0` 构建的模型编写的。如果你的 llamafile 显示的是不同的版本,并且下面的某些步骤无法按预期运行,请在 llamafile 问题跟踪器上提交问题。
使用 llamafile 最简单的方式是通过其内置的聊天界面。在终端中运行
./TinyLlama-1.1B-Chat-v1.0.Q5_K_M.llamafile你的浏览器应该会自动打开并显示一个聊天界面。(如果没有,只需打开浏览器并访问 http://localhost:8080)。聊天结束后,回到终端并按 `Control-C` 关闭 llamafile。如果你是在 notebook 中运行这些命令,只需中断 notebook 内核即可停止 llamafile。
在本教程的剩余部分,我们将使用 llamafile 的内置推理服务器,而不是浏览器界面。llamafile 的服务器提供了一个 REST API,用于通过 HTTP 与 TinyLlama LLM 交互。完整的服务器 API 文档可在此处获取。要以服务器模式启动 llamafile,请运行:
./TinyLlama-1.1B-Chat-v1.0.Q5_K_M.llamafile --server --nobrowser --embedding总结:下载并运行 llamafile
# 1. Download the llamafile-ized model
wget https://huggingface.co/Mozilla/TinyLlama-1.1B-Chat-v1.0-llamafile/resolve/main/TinyLlama-1.1B-Chat-v1.0.F16.llamafile
# 2. Make it executable (you only need to do this once)
chmod +x TinyLlama-1.1B-Chat-v1.0.Q5_K_M.llamafile
# 3. Run in server mode
./TinyLlama-1.1B-Chat-v1.0.Q5_K_M.llamafile --server --nobrowser --embedding使用 LlamaIndex 和 llamafile 构建研究助手
现在,我们将展示如何使用 LlamaIndex 和你的 llamafile 构建一个研究助手,帮助你了解某个感兴趣的主题——在这篇文章中,我们选择了信鸽。我们将展示如何准备数据、索引到向量存储,然后进行查询。
在本地运行 LLM 的好处之一是隐私。你可以混合使用“公共数据”(如维基百科页面)和“私有数据”,而不必担心与第三方共享数据。私有数据可以包括例如你对某个主题的私人笔记或机密内容的 PDF。只要你使用本地 LLM(和本地向量存储),就不必担心数据泄露。下面,我们将展示如何结合这两种类型的数据。我们的向量存储将包括维基百科页面、一本关于照顾信鸽的陆军手册,以及我们在阅读这个主题时记录的一些简短笔记。
要开始,请下载我们的示例数据:
mkdir data
# Download 'The Homing Pigeon' manual from Project Gutenberg
wget https://www.gutenberg.org/cache/epub/55084/pg55084.txt -O data/The_Homing_Pigeon.txt
# Download some notes on homing pigeons
wget https://gist.githubusercontent.com/k8si/edf5a7ca2cc3bef7dd3d3e2ca42812de/raw/24955ee9df819e21975b1dd817938c1bfe955634/homing_pigeon_notes.md -O data/homing_pigeon_notes.md接下来,我们需要安装 LlamaIndex 及其一些集成:
# Install llama-index
pip install llama-index-core
# Install llamafile integrations and SimpleWebPageReader
pip install llama-index-embeddings-llamafile llama-index-llms-llamafile llama-index-readers-web启动你的 llamafile 服务器并配置 LlamaIndex
在这个例子中,我们将使用同一个 llamafile 来生成将被索引到向量存储中的嵌入,并作为稍后回答查询的 LLM。(但是,你完全可以用一个 llamafile 处理嵌入,用另一个单独的 llamafile 处理 LLM 功能——你只需要在不同的端口上启动 llamafile 服务器。)
要启动 llamafile 服务器,打开终端并运行:
./TinyLlama-1.1B-Chat-v1.0.Q5_K_M.llamafile --server --nobrowser --embedding --port 8080现在,我们将配置 LlamaIndex 以使用这个 llamafile:
# Configure LlamaIndex
from llama_index.core import Settings
from llama_index.embeddings.llamafile import LlamafileEmbedding
from llama_index.llms.llamafile import Llamafile
from llama_index.core.node_parser import SentenceSplitter
Settings.embed_model = LlamafileEmbedding(base_url="http://localhost:8080")
Settings.llm = Llamafile(
base_url="http://localhost:8080",
temperature=0,
seed=0
)
# Also set up a sentence splitter to ensure texts are broken into semantically-meaningful chunks (sentences) that don't take up the model's entire
# context window (2048 tokens). Since these chunks will be added to LLM prompts as part of the RAG process, we want to leave plenty of space for both
# the system prompt and the user's actual question.
Settings.transformations = [
SentenceSplitter(
chunk_size=256,
chunk_overlap=5
)
]准备数据并构建向量存储
现在,我们将加载数据并对其进行索引。
# Load local data
from llama_index.core import SimpleDirectoryReader
local_doc_reader = SimpleDirectoryReader(input_dir='./data')
docs = local_doc_reader.load_data(show_progress=True)
# We'll load some Wikipedia pages as well
from llama_index.readers.web import SimpleWebPageReader
urls = [
'https://en.wikipedia.org/wiki/Homing_pigeon',
'https://en.wikipedia.org/wiki/Magnetoreception',
]
web_reader = SimpleWebPageReader(html_to_text=True)
docs.extend(web_reader.load_data(urls))
# Build the index
from llama_index.core import VectorStoreIndex
index = VectorStoreIndex.from_documents(
docs,
show_progress=True,
)
# Save the index
index.storage_context.persist(persist_dir="./storage")查询你的研究助手
最后,我们准备就信鸽提出一些问题。
query_engine = index.as_query_engine()
print(query_engine.query("What were homing pigeons used for?")) Homing pigeons were used for a variety of purposes, including military reconnaissance, communication, and transportation. They were also used for scientific research, such as studying the behavior of birds in flight and their migration patterns. In addition, they were used for religious ceremonies and as a symbol of devotion and loyalty. Overall, homing pigeons played an important role in the history of aviation and were a symbol of the human desire for communication and connection.print(query_engine.query("When were homing pigeons first used?"))The context information provided in the given context is that homing pigeons were first used in the 19th century. However, prior knowledge would suggest that homing pigeons have been used for navigation and communication for centuries.结论
在这篇文章中,我们展示了如何通过 llamafile 下载并在本地设置运行 LLM。然后,我们展示了如何使用 LlamaIndex 将此 LLM 与 LlamaIndex 结合,构建一个简单的基于 RAG 的研究助手,用于学习信鸽知识。你的助手完全在本地运行:你无需支付 API 调用费用,也无需将数据发送给第三方。
作为下一步,你可以尝试使用更好的模型(如 Mistral-7B-Instruct)运行上述示例。你也可以尝试为不同主题(如“半导体”或“如何烤面包”)构建研究助手。
要了解有关 llamafile 的更多信息,请查看 GitHub 上的项目,阅读这篇关于使用 LLM 的 bash 单行命令的博客文章,或在 Discord 上向社区问好。
来源:LlamaIndex:产品、工程与评测 · llamaindex.ai