跳到正文
Every:最新文章·· 3 小时前AI 评分60

Every 发布开源模型入门指南,教你怎么判断哪些任务可以不用闭源前沿模型

Getting Started With Open Models

AI 导读

Every 发布开源模型实用指南,讲解如何评估需求、选择模型与托管方式,将开放模型用于编码、创意和日常自动化等工作。指南指出当下最强开源模型多来自中国实验室,并强调按任务难度区分模型选择,同时附上一个用于审计现有 AI 用法、筛选可迁移到廉价模型任务的提示词。全文为付费订阅内容。

正文

Skip to content

Image 1Image 2

Sign inSubscribe

Image 3: Every

Getting Started With O pen Models

A practical guide to commodity intelligence, owning your AI stack, and deciding when to still reach for the frontier

Image 4: Kai ZauKai ZauImage 5: Opus 4.6Opus 4.6Image 6: GLM 5.1GLM 5.1Image 7: GLM 5.2GLM 5.2

Not long ago, the best frontier models were as unreliable as they were impressive. They could explain quantum mechanics and write decent code, but couldn’t resist fabricating quotes or miscounting the Rs in “strawberry.” Anyone doing serious work used the strongest model they could access.

Since then, the best models have raced ahead in specialized areas like research math, advanced cybersecurity, and scientific discovery—far beyond what most of us need for everyday work. Meanwhile, they’ve grown more expensive to use and more locked down.

But the same progress has also turned AI models into commodities. Capabilities that once required a frontier model are now available in cheaper, widely available models. For routine tasks, today’s open models are competitive on both price and quality, and companies have begun to use them alongside frontier models from OpenAI and Anthropic.

It doesn’t take a genius to summarize meeting notes—though it does take one to discover new antibiotics. Most work falls somewhere in between, and the practical question to ask is, “Which model is good enough for my task?”

I’ve been tinkering with open models since Meta released Llama in 2023. Back then, open models were research curiosities. Today, they handle a growing share of my work, from professional coding to creative projects to automating digital chores.

The cost savings are nice, but the bigger prize is independence. When an API goes down, you have a backup. When a model vanishes from the menu, you have other options. This guide walks through the open model stack one layer at a time. It shows you how to assess your needs, choose a model and host, and put it to work.

Read with ClaudeRead with ChatGPTCopy for agent

The AI stack you can own

When you use ChatGPT or Claude, the app connects to a model running on the company’s servers. The company also stores your conversation history so you can pick up where you left off across devices. Even if you install the app on your computer, the model still runs remotely. The model itself is just a large file containing billions of numbers, called “weights,” that encode the patterns it learned in training.

U.S. frontier labs keep their weights secret. Open-weight models make those files available to download and run yourself, subject to their licenses. Some—but not all—open-weight models are also open source: they share training code and information about the training data, with broad rights to use, modify, and redistribute them.

Model weights are just the start. Every other layer of the product—from hosting to tools to interface—also has open alternatives.

You can decide how much of each layer to run yourself. Paying a provider saves you setup and maintenance work. Running it yourself gives you more say over where your data goes and how the system works, but you become responsible for keeping it running. It’s not all or nothing, and different tasks lend themselves to different setups.

Chinese open models

Most of the strongest open models today come from Chinese labs. U.S. and European labs release open models too, but not as many, and the models they offer are generally not as capable.

Giving models away isn’t charity. Free has long been a competitive business strategy. Google made Android free to phone makers, helping secure a place for its search engine on mobile devices. Chinese labs publish open models for similar reasons: marketing, tapping a global research community, and pressuring U.S. rivals whose business depends on renting out closed models.

If the origin concerns you, consider that Cursor fine-tuned Kimi K2.5 into Composer 2, the model behind its coding agent, and Perplexity has long offered open models alongside frontier ones. Using a model from a Chinese lab doesn’t require using the lab’s app or sending your data to China. A model is a file of numbers. It’s auditable by independent researchers, and it can run on a host you trust.

Negotiating ‘good enough’

Whether a model is good enough depends on what you ask it to do. A task gets harder when the instructions are vague, the consequences of a mistake are serious, or the model has to make many decisions without supervision. The same model can be reliable for one job and inadequate for another.

The strongest models are better at interpreting unclear requests and working through long, complicated problems. Use them when the work is hard to define in advance, takes many steps, or would be costly to get wrong.

Smaller models can handle clear, repeatable work that is easy to verify. Good instructions, relevant reference material, and checks along the way can make them reliable on harder tasks.

Start with the task. Support-ticket triage is routine and mistakes are easy to spot, making it a good candidate for an open model. A one-off payment-system refactor is complicated and costly to get wrong, so use the strongest model you can afford.

The AI you already use may know enough about your work to spot good candidates. This prompt asks it to use any memory or history it has, then fill in the gaps with a few questions.

Prompt Copy

Help me identify work I currently give to frontier AI that may be simple enough for cheaper, faster, commodity AI. Use what you already know about my projects and habits, then interview me to fill the gaps.

For background on open-weight models and the framework used here, see the full guide:

https://every.to/guides/getting-started-with-open-models

The goal is a shortlist of three small experiments, not a complete automation strategy. Prioritize work I already do with AI. You may include an adjacent task that is not currently AI-assisted when it is an unusually strong fit.

Commodity intelligence includes cheaper or open-weight models, whether remote hosted or local. Do not assume I am technical. Keep the conversation focused on what I am trying to accomplish, what goes in, what useful work comes out, and how I would know whether the result is good enough. Leave models, APIs, code, and automation platforms for a later conversation.

How to conduct the audit

  1. Start with evidence you can actually access: our conversation history, saved memories, project files, recurring requests, or other available context. Never imply that you can see history or systems you cannot access. Do not ask me to assemble an inventory before checking what you already know.
  2. Form a tentative picture of my recurring work, then interview me to correct and complete it. Ask one or two focused questions at a time, adapting each question to what you have learned. Do not give me a long questionnaire.
  3. Look beyond job titles and broad responsibilities. Identify concrete tasks with recognizable beginnings and endings. “Manage marketing” is too broad. “Turn the weekly campaign metrics into Monday’s update” is a task.
  4. Use specific examples from my past work when available. Ask for a recent example when an attractive candidate remains abstract. Quietly correct obvious transcription errors and distinguish a true routine from something that merely happened twice in similar language.
  5. Explore a typical day, week, and month only where the available history leaves gaps. Pay particular attention to work I delegate to AI repeatedly, work I perform from a template, and chores I postpone because they are tedious or expensive.
  6. Notice where each task already lives. Look for an established trigger, source, interface, system of record, or destination: a recurring meeting, inbox, folder, spreadsheet, project tool, chat habit, or regular document. Favor experiments that can join an existing routine over ideas that require me to adopt a new one.
  7. Keep discovery separate from implementation. Do not recommend tools, models, vendors, integrations, or technical architecture during the audit unless I explicitly ask.

How to judge candidates

A strong candidate usually has several of these qualities:

  • It recurs often enough that learning or setup can pay off.
  • Each run begins with known or recognizable inputs.
  • The desired output can be described concretely.
  • Past examples show what acceptable work looks like.
  • Success can be checked without relying entirely on taste.
  • The task is narrow enough to complete in one bounded pass.
  • A weak result is cheap to catch, discard, correct, or rerun.
  • It can run in the background without keeping me waiting.
  • A frontier model already handles it comfortably and consistently.
  • It connects to a habit or system I already use.

Treat these as judgment criteria, not a rigid numerical score. Frequency alone does not make a task suitable. For each candidate, also test:

  • Scope: How much ambiguity, judgment, context, and multi-step planning does it require?
  • Stakes: What happens if the result is subtly wrong? Can it send, spend, publish, delete, advise, or otherwise create consequences before review?
  • Verifiability: Can I recognize success cheaply, or would checking the work take as long as doing it?
  • Economics: Does it happen often enough to justify changing the routine? If my frontier use is already covered by a flat subscription, independence, privacy, capacity, or flexibility may matter more than token savings.

Do not force an entire task onto commodity intelligence when only part of it fits. Look for a safer boundary: commodity intelligence can extract, sort, summarize, format, or prepare a first pass while a person or frontier model retains the ambiguous judgment. Prefer a smaller useful experiment to an impressive autonomous system.

Keep frontier intelligence for work that is one-off, poorly defined, hard to verify, taste-dependent, consequential, or cheap only when it succeeds on the first try. When a candidate does not survive scrutiny, say so at that point in the conversation and explain why. Do not save rejected tasks for a required section in the final answer.

Present the shortlist in layers

Once you understand enough of my work, present the result in four layers. Give me the recommendation before the audit trail.

What I noticed

Open with a conversational summary of the pattern you found in two to four sentences. Focus on what makes parts of my work promising for commodity intelligence, not on recapping your interview process.

Three experiments worth trying

Present the best three candidates in priority order as a numbered list. Give each a plain-language name followed by two or three sentences. Tell me what the experiment would accomplish, why this task appears suitable, and where human or frontier judgment should remain. Write this as advice to me, not as a form or scorecard.

Prefer a short, credible list over filling the quota with weak ideas. If fewer than three tasks genuinely fit, return fewer and explain what evidence is missing.

Experiment details

Follow the summary with compact supporting details for each candidate. Use the same names and order so the sections are easy to connect. Include:

  • Current routine and foothold: When it happens, how I handle it now, and the existing habit or system it can join.
  • Input to output: What a typical run starts with and the concrete result I need.
  • Safe boundary: What should remain under human or frontier judgment, including why the broader task would be unsuitable when relevant.
  • Smallest trial: A reversible comparison using a few past or parallel examples, without changing the live routine.
  • Pass condition: The observable result that would make the experiment worth repeating.
  • Main doubt: The likeliest reason it may fail or prove uneconomical.

Keep these details concise and avoid repeating the recommendation above. They should support it, not bury it.

Where I would start

Recommend one candidate to try first and explain the choice in two or three sentences. Favor easy learning and low downside over maximum theoretical savings. End by asking whether I want to refine that experiment or turn it into a lightweight evaluation plan. Do not proceed into implementation until I choose.

This guide is for paid subscribers

Upgrade your membership to continue reading the full guide.

Subscribe to continue

Benchmarks: public and private

Guide workflow

Cost per task, not per token

Guide workflow

Step 1: Choose your model

Guide workflow

Extra large: Frontier rival

Guide workflow

Large: Capable agent

Guide workflow

Medium: Competent worker

Guide workflow

Small: Junior worker

Guide workflow

Tiny: Pocket utility

Guide workflow

Step 2: Choose your hosting

Guide workflow

Hosted inference providers

Guide workflow

Self-hosting

Guide workflow

Step 3: Integrating your work

Guide workflow

Existing apps

Guide workflow

Open alternatives

Guide workflow

Automated workflows

Guide workflow

Intelligence as a commodity

Guide workflow

Image 8

See how good you are at using AI—and learn how to get better

Sign in and get your AI assessment prompt.

Image 9Continue with GoogleView all login options

Image 10: Every

What Comes N ext

New ideas to help you build the future—in your inbox, every day.

Email address

Image 11

Subscribe

Do Not Sell or Share My Personal Information

This site is protected by reCAPTCHA and the GooglePrivacy Policy andTerms of Service apply.

©2026 Every Media, Inc.

AboutCareersHelp centerPrivacy preferencesAdvertise with usThe teamPricingFAQTermsSite map

XImage 12: Arrow OutLinkedInImage 13: Arrow OutYouTubeImage 14: Arrow Out

©2026 Every Media, Inc.

We use analytics and advertising tools by default. You can update this anytime.

Privacy Preferences Do Not Sell or Share

Privacy Preferences

x

Manage optional tracking categories. Necessary cookies stay on so the site can function.

Analytics- [x] Advertising and sharing- [x]

Cancel Save Preferences

Getting Started With Open Models - Every

来源:Every:最新文章 · every.to