跳到正文
原文
MarkTechPost(RSS)· Asif Razzaq·· 3 小时前AI 评分66

Perplexity 发布 Rust 检索引擎 Photon,p99 延迟从约 800 ms 降至约 65 ms

Perplexity Introduces Photon: A Rust-Based Retrieval Engine That Cuts p99 Latency From 800 ms to 65 ms

AI 导读

Perplexity 发布自研 Rust 检索与排序引擎 Photon,已承接全部生产流量,p99 检索排序延迟从约 800 ms 降至约 65 ms,服务机器减少约 20%。

正文

Perplexity has released Photon, an in-house retrieval and ranking engine written in Rust. It replaces an open-source engine Perplexity had forked for its AI-native search stack. Photon now handles retrieval and ranking for all production traffic. It also powers a new Fast Search mode in the Perplexity Search API. Perplexity reports single-call latency of 160 ms at p50 and 230 ms at p95.

Is it deployable? Yes, as a hosted API. Set search_type: "fast" on POST /search and pay $1 per 1,000 requests. Photon itself is not open source, so the engine cannot be self-hosted.

Why Perplexity Replaced its Old Engine

The old engine hit 3 limits as the index grew:

  • Tail latency: Production p99 sat near 800 ms. The dataset exceeded RAM, so mlock was not an option. Cold reads triggered major page faults that stalled queries.
  • Merge spikes: During disk index fusion, p99 climbed to about 1.2 s for 10 to 15 minutes.
  • Slow recovery: Deploying and syncing an extra cluster could take more than a week. Recovery also raised the share of partial responses.

Perplexity team concluded that building from scratch was simpler and cheaper than maintaining its fork.

How Photon Works

A load balancer routes each request to a Photon broker. The broker fans out to a shard group and watches for timeouts. Each shard runs retrieval, initial ranking, and second-stage ranking. The broker then merges candidates and fetches key document fields.

  • Adaptive posting lists: Short lists sit inline within a single page. Longer lists split into blocks of fixed document ID ranges. Sparse blocks store sorted offset arrays and use galloping search. Dense blocks use bitmaps, so membership becomes a single bit lookup.
  • Budgeted traversal: A WAND-like algorithm splits lists into driving lists and probe lists. Cheap presence checks bound each candidate’s maximum score first. Exact term frequencies are read only when a candidate can clear the threshold.
  • Docblob records: Each document gets a compact record of frequencies, field masks, and positions. Terms use Elias-Fano encoding, so ranking decodes only the matched terms. Ranking a candidate needs just 1 lookup per document.
  • Batched async reads: Record offsets are known upfront, so disk reads go out in batches through io_uring. The cache checks the whole batch first. Readers take no locks, and eviction uses CLOCK instead of a shared LRU list.
  • Separate build and serve: Indexers build versioned shard indexes from YTsaurus tables on dedicated nodes. A controller rotates serving groups one at a time and warms caches with replayed search-log queries.

A full web index now builds in a single-digit number of hours.

Interactive Explainer: Inside Photon

Production Results

  • p99 retrieval and ranking latency fell from about 800 ms to about 65 ms. This covers Photon’s stages only.
  • Photon runs on about 20% fewer serving machines than the old content nodes.
  • It stores about 2.5x as much data per document, which Perplexity used to improve ranking quality.
  • Pinning the same dataset with mlock would need an estimated 4.6x the resident memory Photon uses today.
  • Index version switches no longer cause latency spikes.

Fast Search: Speed and Cost for Agents

Fast Search pairs Photon with lighter ranking tuned for agentic workflows. Perplexity tested it on 6 benchmarks: WideSearch, BrowseComp, DSQA, FRAMES, SEAL-0, and SEAL-Hard. Across 3,554 tasks, Fast scored 64.3% at $59.73 in estimated model-plus-search cost. The default preset scored 64.0% at $187.60, so Fast was about 68% cheaper.

The trade-off shows up in broader search quality. On internal long-tail benchmarks, relevance (DCG) fell from 2.45 to 2.21. Answer availability dropped from 0.596 to 0.567, a loss of 2.9 percentage points. Perplexity recommends Fast for day-to-day agent loops and the default for hard, ambiguous queries.

Copy CodeCopiedUse a different Browser

curl -X POST 'https://api.perplexity.ai/search' \
  -H "Authorization: Bearer $PERPLEXITY_API_KEY" \
  -H 'Content-Type: application/json' \
  -d '{"query": "latest stable Rust release", "search_type": "fast", "max_results": 5}'

On Python SDK 0.43.4 and 0.43.5, pass extra_body={"search_type": "fast"} per the docs.

Fast Search vs Closest Competitors

FeaturePerplexity Fast SearchExa InstantParallel Search TurboTavily ultra-fast
Request parametersearch_type: "fast"type: "instant"mode: "turbo"search_depth: "ultra-fast"
Vendor-reported latency160 ms p50, 230 ms p95 (blog)~250 ms typical (docs); sub-200 ms at launch~200 ms (docs)No figure published; lowest-latency depth (docs)
List price per 1K requests$1 (pricing)$4 for up to 10 results (pricing)$1 (docs)1 credit: $8 pay-as-you-go, $5 to $7.50 on plans (credits)
Results per request1 to 2010 in base price, $1 per 1K per extra resultNot specifiedNot specified
Known limitsLower relevance than default presetExtra results billed separatelyEnglish and Japanese queries onlyLower relevance than other depths
LaunchedSep 24, 2026Feb 12, 2026Jul 13, 2026 (blog)Jan 5, 2026 (blog)

All latency figures are vendor-reported under different setups, so they are not like-for-like.

Key Takeaways

  • Photon is Perplexity’s Rust retrieval and ranking engine, now serving all production traffic.
  • Production p99 latency dropped from about 800 ms to about 65 ms.
  • Fast Search reports 160 ms p50 and 230 ms p95 at $1 per 1,000 requests.
  • Fast cut estimated agent task cost by about 68% at comparable task quality.
  • It trades some retrieval relevance, so keep the default preset for hard queries.

Check out the technical details and Fast Search docs. All credit goes to the researcher of this project. Also, feel free to follow us on Twitter and don’t forget to join our 150k+ML SubReddit and Subscribe to our Newsletter. Wait! are you on telegram? now you can join us on telegram as well.

Need to partner with us for promoting your GitHub Repo OR Hugging Face Page OR Product Release OR Webinar etc.? Connect with us

The post Perplexity Introduces Photon: A Rust-Based Retrieval Engine That Cuts p99 Latency From 800 ms to 65 ms appeared first on MarkTechPost.

来源:MarkTechPost(RSS) · marktechpost.com