跳到正文
elvis· @omarsar0 · X·· 3 小时前AI 评分74
AI 导读

Elvis Saravia 回应 Karpathy 关于理解 LLM 输出的帖子,称类似做法自己已实践一年多,并分享他正在搭建的通用人机协作界面。其当前方案将高级智能体连接到类 Notion 的灵活页面,支持嵌入工件、文本、视频和可视化讲解,智能体按任务输出这类工件而非默认文本回复,他通过评论和建议直接协作,覆盖代码审查、笔记、写作、研究和原型设计等场景。

正文

50K likes on this!?

Most of what Karpthy has shared is things we have been doing for over a year.

No surprises, as many replies are highlighting.

But I still think this raises an interesting discussion.

Is there a universal interface/langauge that allows more seamless communication between agents and humans?

I believe there is. And I think that is what @karpathy is suggesting we think about.

Right now, ideas are all scattered.

As some of you know, I have been early on LLM Wikis, Artifacts, AI-generated visual explainers, dynamic interactives, etc.

Over the past couple of months, I have been experimenting with a new universal interface for seamless human-agent collaboration. So I have a lot to share, but it's still too early, and I frankly haven't found the perfect solution.

In short, my current optimal setup (see a simple snapshot attached) ties my high-level agent to highly flexible Notion-like pages that support embedding artifacts, text, videos, visual explainers, and just about anything an LLM can generate. And if it doesn't support it yet, I ask it to build it on the fly. It's crazy sometimes, the features it builds, but this is a co-evolved interface that satisfies both my agent's and my needs.

The interface supports code review, note-taking, writing, research, prototyping, mock designing, and so on.

In fact, I have now wired my agents to intelligently spit out these artifacts (for lack of a better word), depending on the task/instruction, instead of the default text responses, which are frankly overwhelming and hard to read at times. In other words, as I keep building this "universal interface," it feels less and less like I am chatting with an agent; I'm collaborating with it directly through the artifacts (through comments and suggestions). It allows me to move faster. I can consume and review faster. I am better organized. It feels like a language both the agent and I can speak fluently, so there is very little friction.

Listen, this is all a personal experience. I don't think everyone will use interfaces like this. But it's been interesting to see Artifacts gain wide adoption, and now dots have Spaces that look a little like what I built here several months ago.

More to share soon.

引用Andrej Karpathy@karpathy
We'll be spending a lot more time trying to understand the outputs of language models. A few thoughts, tips & tricks: Writing. Something I've had success with: Ask your LLM to explain something in ASD-STE100, it's a controlled language specification originally developed for aerospace maintenance documentation. LLMs well-versed in this language and it comes with heavy constraints on clean writing style that I often find a lot more readable. Sometimes I've tried to soften it a bit e.g. ask for "80% of the way to ASD-STE100" because the spec is quite stringent. But even better: Diagrams / images. Instead of writing, ask your LLM to create a diagram. These can be a lot easier to process, parse, and understand. But even better: Web pages. Ask for output "in HTML" to get a beautiful, interactive webpage. LLMs are getting really good at frontend and can create beautiful experiences, animations, etc. But even better: Explainer videos. The output format I am most bullish on is fully custom / bespoke explainer videos generated on any arbitrary topic. Experiment with things like "Create a 3b1b style video explainer on X. Use my ElevenLabs API key for audio narration". (you'd need an API key for the latter or you can ask your LLM to find you decent free alternatives that use your local compute). This is actually starting to work! In summary: - As LLMs get better, they will do more and more of the legwork autonomously, and a lot more of our work will rise up the abstractions into oversight and understanding. - Luckily, LLMs can help here too because as intelligence and code are increasingly abundant, you can ask for large, custom, discardable software artifacts (e.g. web apps, video explainers) that would have never made sense to create before. Push the boundaries here and you'll be surprised.
在 X 查看被引用的帖子

来源:elvis · x.com