跳到正文
原文
MarkTechPost(RSS)· Asif Razzaq·· 3 小时前AI 评分53

Perplexity 发布上下文嵌入模型 pplx-embed-v2-context-9b-preview,可同时检索答案与支撑证据

Perplexity Releases pplx-embed-v2-context-9b-preview: A Contextual Embedding Model That Retrieves Answers and Their Supporting Evidence

AI 导读

Perplexity Research 与 turbopuffer 发布上下文嵌入模型 pplx-embed-v2-context-9b-preview,权重以 MIT 许可托管在 Hugging Face,可自托管部署,尚未上线 Perplexity API。

正文

Perplexity Research and turbopuffer have released pplx-embed-v2-context-9b-preview, a contextual embedding model for RAG pipelines. Each chunk is embedded with the full document in view. The real change is the training signal. The model learns to retrieve the answer along with the context needed to verify it, not one ‘gold passage.’ P

Is it deployable? Yes, as a self-hosted preview. Weights are on Hugging Face under the MIT license. Loading requires transformers>=5.4.0 with trust_remote_code=True. It is not yet on the Perplexity API. The model card notes that weights and interface may change without backward compatibility.

Why the gold passage falls short

RAG systems split long documents into chunks. A chunk often depends on an entity, heading, or definition stated elsewhere. Contextual models address this with late chunking. The document is encoded in one pass, then pooled per chunk.

Training, however, usually marks one gold chunk per query. Every other chunk becomes a negative, including the sentences that make the answer checkable. Perplexity lists 3 more problems. Binary labels give a coarse signal. LLM annotation cost grows linearly with dataset size. Labels are also tied to one chunking strategy.

How the training works

The teacher is Perplexity’s query-aware context compression model. It reads the query and document together and scores every token.

  • Chunk relevance: the mean of the top n token scores inside each chunk.
  • Soft target: a temperature-scaled softmax over chunks in the positive document. Chunks in other documents get zero.
  • Distillation loss: forward KL divergence between teacher and student distributions.
  • Document loss: InfoNCE, where a document scores as its best chunk, inspired by ColBERT’s MaxSim.

Each batch samples a random chunking strategy. Chunks are separated by a learned <|chunk_sep|> token and mean-pooled. The teacher runs only during training, so inference adds no latency or storage.

The model starts from an in-house 9B ColBERT retrieval model. A linear projection outputs 2048 dimensions. Matryoshka training also supports 1024 dimensions. Quantization-aware training enables native int8 embeddings. The release is a model soup of several checkpoints. Training used roughly 430 datasets covering over 50 languages, with no ConTEB data.

Interactive explainer

ctx

How pplx-embed-v2-context retrieves answers and their evidence

An interactive walkthrough of Perplexity’s contextual embedding preview: why isolated chunks fail, how teacher distillation replaces the single gold passage, and what the reported numbers mean.

Three lease files share the sentence “Monthly rent is …”. Only one belongs to 5 Park Avenue. Switch modes and run retrieval.

QUERYWhen does 5 Park Avenue’s lease end and what is the current rent?

Pick a mode and press Run retrieval.

Illustrative example modeled on the lease scenario in Perplexity’s post. Scores are for explanation only, not model outputs.

A context compression model acts as teacher. It scores every token for the query. Those scores are pooled per chunk (mean of the top n tokens) and turned into a soft target, instead of a one-hot gold label.

Chunk boundaries:

Temperature0.25

1 Teacher scores tokens2 Top-n mean per chunk3 Softmax target4 Student matches via KL

Gold-chunk label (one-hot)

Teacher target (soft, from token scores)

Change the boundaries: the same token scores re-aggregate without re-annotation. That is the “flexible chunk boundaries” property Perplexity describes. Token scores here are illustrative.

context-bench (2,099 queries, 38,894 documents, 2,458,072 sentence chunks, exhaustive ranking). Numbers below are as reported by Perplexity at K = 10.

pplx-embed-v2-context-9b-previewvoyage-context-4 (derived from reported gap)

Voyage values are computed as Perplexity’s figure minus the stated gap (14.4 and 5.0 points). Other Voyage metrics appear only in Perplexity’s chart and are not shown here.

Contextual embeddings store one vector per chunk, same as a normal chunk index. Cost depends on vector size. Perplexity reports that 1024-dim int8 (1 KB) slightly exceeds voyage-context-4 at 2048-dim float32 (8 KB) on its chunk-retrieval suite.

Chunks

0

vector storage (vectors only, not full index)

Bytes per vector

–

vs 2048-d float32

–

Chunk-size sensitivity (64 to 512 tokens)

81.0% to 79.9%

Bytes = dimensions x bytes per value. Sensitivity is mean nDCG@10 across 74 MTEB tasks, as reported by Perplexity.

Sources: Perplexity Research · Model card Built by Marktechpost

context-bench: a new benchmark

context-bench is built and privately held by turbopuffer to limit training contamination. It holds 2,099 queries over 38,894 documents in 21 domains. Sentence chunking yields 2,458,072 chunks. Median target document length is roughly 6,100 tokens. Queries test 12 contextual capabilities, from pronoun resolution to table structure. Perplexity says the model was submitted blind.

Metrics are Document@K, Answer@K, Evidence Recall@K, and All-Evidence@K. Every model is ranked exhaustively against all chunks, so index settings play no role.

Results

  • context-bench at K = 10: 45.5% answer recall, 40.6% evidence recall, 31.1% all-evidence recall.
  • Document recall: 15.2% at K = 1 and 61.6% at K = 10.
  • vs voyage-context-4: ahead by 14.4 points on answer recall and 5.0 on evidence recall at K = 10.
  • ConTEB: highest average nDCG@10 among models shown. pplx-embed-context-v1-4B wins NarrativeQA, and Nemotron-3-Embed-8B wins COVID-QA.
  • General retrieval: best average on query-to-chunk tasks. Slightly behind voyage-context-4 on query-to-document.
  • Storage: 1024-dim int8 (1 KB per vector) slightly beats voyage-context-4 at 2048-dim float32 (8 KB).
  • Chunk size: average score moves from 81.0% to 79.9% between 64 and 512 tokens.

Comparison with the closest competitors

Featurepplx-embed-v2-context-9b-previewvoyage-context-4pplx-embed-context-v1-4BNemotron-3-Embed-8B
Chunk embeddingsContextualContextualContextualIndependent per chunk
AccessOpen weights, MITHosted API (Voyage, MongoDB Atlas)Open weights, MIT; Perplexity APIOpen weights, OpenMDW-1.1
Parameters9B per blog (Hugging Face lists 8B)Not disclosed (MoE backbone)4BAbout 8B
Dimensions2048, 10242048, 1024, 512, 2562560 (Matryoshka)4096, sliceable
Quantized outputNative int8int8, uint8, binary, ubinaryint8, binaryFloat
ContextEvaluated up to 32,768 tokens32K per request; 120K with auto-chunking32K32,768
Auto-chunkingNoYesNoNo
PriceSelf-host$0.12 per 1M tokens; first 200M freeSelf-host or APISelf-host

Sources: Voyage docs, model cards linked above. Checked September 30, 2026.

Key Takeaways

  • A token-level teacher replaces the single gold-chunk label.
  • The model retrieves answers plus supporting evidence in one chunk index.
  • context-bench Answer@10 is 45.5%, 14.4 points above voyage-context-4.
  • Open MIT weights ship today; Perplexity API access is still pending.

Check out the technical details and model weights. All credit goes to the researcher of this project. Also, feel free to follow us on Twitter and don’t forget to join our 150k+ML SubReddit and Subscribe to our Newsletter. Wait! are you on telegram? now you can join us on telegram as well.

Need to partner with us for promoting your GitHub Repo OR Hugging Face Page OR Product Release OR Webinar etc.? Connect with us

The post Perplexity Releases pplx-embed-v2-context-9b-preview: A Contextual Embedding Model That Retrieves Answers and Their Supporting Evidence appeared first on MarkTechPost.

来源:MarkTechPost(RSS) · marktechpost.com