今日已更新 446 条资讯 | 累计 23583 条内容
关于我们

今日精选

HOT

最新资讯

共 23583 篇
第 261/1180 页
AI 资讯 Dev.to

Anthropic Just Admitted MCP Has a Context Problem

Anthropic Built a Fix That Proves the Problem 🔍 Right, so Anthropic dropped Tool Search on November 24th alongside Claude Opus 4.5, and I need you to sit with the implications for a second because they're brutal for MCP. Tool Search lets you mark tools as defer_loading: true . When you do that, the tool's schema never enters your context window until Claude actually needs it. Claude gets a name and nothing else. When a task comes in that requires the tool, Claude calls the Tool Search tool, pulls the schema, and only then does it have the full definition loaded. Lazy loading for AI tools. Sounds boring. It's not. Here's the number that should make you spit your tea out: accuracy went from 49% to 74% on Opus 4. On Opus 4.5, it climbed from 79.5% to 88.1%. Same model. Same tools. The only difference is not loading the tool definitions upfront . I'll say it again. The mere presence of tool schemas in context was tanking accuracy by 25 percentage points. Twenty-five. Not because the model got confused by irrelevant tools (though it did). Because the context window was full of JSON nobody was using, and the model was drowning in it. This is an admission dressed up as a feature I've been banging on about MCP's context cost since July. The Playwright MCP server alone burns 12.8k tokens just sitting there. Load four or five MCP servers and you've torched half your context window before asking the model to do anything useful. Tool Search is Anthropic's answer. And the answer is: don't load the tools. Think about that for a second. The company that created MCP built a feature whose entire purpose is avoiding the cost of loading MCP tool definitions. If that isn't an admission that the protocol has a context problem, I don't know what is. They didn't say "we've made MCP more efficient." They said "here's a way to not load MCP schemas until the last possible moment." The fix is the diagnosis. And the timing is proper interesting. Three weeks earlier, on November 5th, mcporter d

Steven Gonsalvez 2026-07-10 23:44 3 原文
AI 资讯 Dev.to

Peter Steinberger Says Just Talk To It, and He's Mostly Right

The Man Behind the Minimalism I met Peter Steinberger at Claude Anonymous in London sometime in June, and the presentation was properly fun. He's got that energy where you can tell he's not performing, he's just genuinely excited about the stuff he's building. Maintains a 300k line TypeScript React ecosystem. Web app, Chrome extension, CLI tool, Tauri desktop client, Expo mobile app. All maintained by one person with AI agents. And his setup is basically: open 3-8 parallel Codex instances, write short prompts (often 1-2 sentences plus a screenshot), let them go. He calls everything else "charade." I've been following his work since and I reckon he's one of the most practical voices in the agentic coding space. No hype. No frameworks for the sake of frameworks. Just a bloke who ships a lot of code with AI and has opinions about how to do it well. His Principles (and Where I Stand) Here's the full list from his talk and blog posts. I agree with most of it. Where I don't, I'll say so. Things I violently agree with: Use tmux to run CLIs persistently. Absolutely. This is foundational. If you're not doing this, start here. Use ast-grep as a pre-commit hook. Proper structural linting that catches things regex can't. Brilliant addition to any agent workflow. Ask for options before executing. I built a whole /interview skill around this. Get the model to present choices rather than guessing. Way better outcomes. Don't reset context. Cold start wastes time + tokens. Agreed, but I'd go further. Don't reset, /handover instead. Summarise what was accomplished, carry it to the next session. Cold start is a waste. Context reset without transfer is nearly as bad. Give examples. Multi-shot prompting is always better than zero-shot. If you want the model to output something specific, show it what that looks like first. Massively improves one-shot accuracy. Use voice input. Absolutely. I started with Superwhisper and it changed how I work. Speaking your intent while pacing around the

Steven Gonsalvez 2026-07-10 23:44 4 原文
AI 资讯 Dev.to

Ultrathink and Build: Weekly Dev Log with AI Tools and Side Projects

What I'm Reading and Watching This Week 📚 [Article] " Deep Work in the Age of AI " by Cal Newport. Proper interesting take on how AI coding tools change our relationship with focused work. Made me rethink a few things about how I structure my own deep work blocks. [Video] "The Art of Code Review" by ThePrimeagen. Some good bits on making code reviews actually useful instead of the usual rubber-stamping faff. [Book] Currently reading "Staff Engineer" by Will Larson. If you're a senior dev wondering what comes next, this is the one. Side Projects and AI Dev Tools I'm Building 🛠️ stevengonsalvez.com 🟢 Active Next.js 15 + MDX blog platform with the byte-sized banter section you're reading right now. Foundation's done, blog section coming together. Next up is deploying to Vercel and sorting the domain. Personal MCP Server 🟡 On Hold Custom MCP server for personal productivity tools. Paused while I wait for Claude Desktop to get its MCP support properly sorted. What Changed This Week 🗞️ Migrated the blog from dev.to-only to a self-hosted Next.js site. Launched this new weekly banter format. Still experimenting with how often I actually want to publish, reckon weekly is about right but we'll see. Developer Tips and Tricks 💡 Quick git alias for better logs: git config --global alias.lg "log --graph --oneline --all --decorate" This creates a beautiful visual git history that's way more readable than the default log. Claude Desktop optimization tip: Keep conversations small and restart often. The message limit resets every 5 hours, but shorter conversations use fewer tokens per message. You can get 2-3x more usage by starting fresh chats instead of continuing long threads. (If you're weighing up Claude Code versus other terminal tools, I did a proper comparison of Claude Code vs Warp AI a while back.)

Steven Gonsalvez 2026-07-10 23:44 3 原文
AI 资讯 Dev.to

Claude Code Skills Just Made Half Your MCP Servers Redundant

Skills are what MCP should have been 📝 Anthropic announced Agent Skills on October 16th, and I reckon this is one of those quiet releases that ends up mattering more than the flashy model drops. A skill is a markdown file with YAML frontmatter. That's it. No server process. No JSON-RPC. No WebSocket connections. No Docker containers. A markdown file. And the clever bit is how it loads. Level 1 is metadata. Name, description, trigger conditions. Costs you about 30 to 100 tokens and it's always in context. Level 2 is instructions. The actual "how to do the thing" content. Under 5,000 tokens. Only gets loaded when the model decides it's relevant to your task. Level 3 is resources. Referenced files, example code, whatever. Only pulled in when the instructions explicitly reference them. So you've got a system where 100 skills can sit in your context at a cost of maybe 3,000 to 10,000 tokens total. The metadata layer alone. The model reads the names and descriptions, figures out which ones matter, and loads only what it needs. Now compare that to MCP. Four or five MCP servers and you're looking at 40,000 to 60,000 tokens of JSON schemas. Loaded upfront. All of them. Whether you need them or not. Sitting in your context window like furniture in a flat you never use, taking up space and making the place harder to navigate. Most MCP servers were never about live data Here's the thing I keep coming back to. What were people actually using MCP for? Some of it was legitimate live data connections. Database queries. API calls. Fetching real-time information the model can't have in its training data. Fair enough. That's a genuine use case and skills can't replace it. But a huge chunk of the MCP ecosystem was procedural knowledge. How to deploy this thing. How to format commits. How to run the test suite. How to interact with Jira. Step-by-step instructions wrapped in a protocol layer and loaded as tool definitions. Tens of thousands of tokens to tell the model "when someone asks

Steven Gonsalvez 2026-07-10 23:44 5 原文
AI 资讯 Dev.to

AI Security Breaches, Vibe Coding Secrets Leak, and OpenAI's $500B Week

AI Security Disasters This Week 🔥 Half a trillion in AI valuations. Eight thousand children's records nicked. Sudo is broken. Everything is on fire and we're shipping vibes. Right. Where to start with this week of AI security breaches and vibe coding gone wrong. A ransomware gang called Radiant (love the branding, lads) hacked a nursery chain called Kido and walked off with personal data on 8,000 children. Children. Not enterprise accounts. Not crypto wallets. Actual kids in actual nurseries. I don't usually get properly miffed about security news because frankly if you're still leaving RDP open to the internet you deserve what's coming. But targeting nurseries? That's a new kind of grim. Meanwhile: OpenAI hit a $500 billion valuation. Half. A. Trillion. We're living in a timeline where AI companies are worth more than most countries' GDP and a nursery chain can't keep toddler data safe. Cool. This is fine. The Hits Keep Coming 💀 It wasn't just the nursery hack. This week was a proper security shambles from top to bottom. Google Ads serving trojans. You search for something legitimate, click an ad at the top of Google, and congratulations, you've just installed malware. Google taking money to distribute malware is chef's kiss levels of ironic. The ad platform that prints money can't vet what it's printing. Fake invoices spreading RATs. Remote Access Trojans shipped via invoice PDFs. Because apparently we still haven't sorted out "don't open random attachments" after twenty-odd years of trying. The sudo exploit (CVE-2025-32463) got added to CISA's Known Exploited Vulnerabilities list. Actively exploited. In the wild. Right now. Sudo. The thing that literally gates root access on every Linux box you've ever touched. If that doesn't make you sweat a bit, you're not paying attention. Tile tracking devices flagged as a stalking risk. The thing you bought to find your keys can apparently be weaponised to find you . Reassuring. UK Co-Op attack costs hit $275 million. Drago

Steven Gonsalvez 2026-07-10 23:43 1 原文
AI 资讯 Dev.to

Multi-Agent Error Cascades: The Double Pendulum Problem Nobody Talks About

Your agents are a chaos machine 🎯 So HumanLayer dropped their Advanced Context Engineering piece last week, and buried in it is this absolute banger of a line: "Bad research -> bad plan -> bad code. A single wrong line in research cascades to widespread errors." That's Dex Horthy describing what happens when you chain AI agents together without checkpoints. And it's the exact same mechanics as a double pendulum. You know the double pendulum, yeah? Simple physics demo. One pendulum hanging from another. The top one swings predictably. The bottom one goes absolutely mental. Tiny changes in the initial swing of the top pendulum produce wildly different trajectories in the bottom one. Chaos theory in action. Looks like it's possessed. Multi-agent AI systems are double pendulums. Your research agent makes a small mistake. Misidentifies which module handles authentication. Gets one import path wrong. Reads an outdated API signature. Tiny error. Barely noticeable in the research output. Your planning agent reads that research and builds a plan around the wrong assumption. The plan isn't obviously wrong. It's coherent. It just starts from a slightly incorrect foundation. The deviation from reality is bigger now but still plausible-looking. Your implementation agent reads that plan and writes hundreds of lines of code. Code that's internally consistent but built on a foundation of sand. The cascade is complete. One wrong line in research became a hundred wrong lines in code. Why more agents makes this worse, not better Here's the bit that does my head in. The whole pitch of multi-agent systems is "more agents = better results." Specialise each agent. Divide the labour. Sounds reasonable. But every agent you add to the chain is another joint in the pendulum. Every handoff is another point where small errors amplify. A three-agent pipeline (research, plan, implement) has two handoff points. A five-agent pipeline has four. Each one is a potential chaos amplification. Sean Moran

Steven Gonsalvez 2026-07-10 23:43 1 原文
AI 资讯 Dev.to

Why Terminal AI Coding Agents Are Beating IDE Extensions

The Headless Takeover 🖥️ IDEs had a good run. Then the terminals learned to think. Here's the thing nobody's saying out loud: terminal-based coding agents are eating headed apps for breakfast. And it's not even close. The Architecture Gap Every IDE-based agent - Cursor, Windsurf, whatever's launching next week - has the same fundamental constraint: IPC . Inter-process communication. The IDE runs here, the AI runs there, and they talk through some protocol layer. Extensions, language servers, message passing. It works, but it's a bottleneck by design. Terminal agents? Direct subprocess execution. No middleware. No protocol translation. You spawn a process, it runs, you get output. Done. 📚 Geek Corner IPC vs Subprocess : IDE extensions typically communicate via JSON-RPC, WebSockets, or custom protocols. Each message serialises, transmits, deserialises. Terminal agents skip all of that - they're just running shell commands as child processes with direct stdin/stdout pipes. The overhead difference is orders of magnitude. The Composability Factor Terminal tools are composable by default . cat file.ts | grep "TODO" | wc -l Thirty years of Unix philosophy baked into every interaction. Pipes. Redirects. Scripts. You can wrap anything - linters, formatters, compilers, test runners, deployment scripts - and the agent just treats them as tools. Try doing that in a GUI. You're clicking buttons. Waiting for panels to load. Fighting with extension APIs that change every release. Terminal agents don't care about your UI framework. They run commands. Commands compose. That's it. The Scaling Play This is where it gets spicy. How do you scale a GUI-based agent? You don't. One human, one screen, one IDE instance. Maybe you run a second window if you're feeling fancy. Terminal agents? Spawn ten of them. Spawn a hundred. Run them in CI. Run them on every PR. Run them overnight while you sleep. They're just processes - orchestrate them however you want. # This is trivial with terminal ag

Steven Gonsalvez 2026-07-10 23:43 1 原文
AI 资讯 Dev.to

GPT-5, Opus 4.1, and Duct-Tape Security: AI's Wildest Week in 2025

The Week That Had Everything 🎪 A 24-year-old lands a $250M AI pay package. Meanwhile, link wrappers are nicking your login credentials. Same industry. Same week. Right, where do I even start with this one. Week 32 was the kind of week where every newsletter hit different. GPT-5 dropped. Opus 4.1 dropped. OpenAI went open-source. Cursor got a CLI and got poisoned. North Korean devs are still out here catfishing hiring managers. And someone, somewhere, is writing a quarter-billion-dollar cheque to a researcher who can't legally rent a car in most US states. Let's crack on. The Money's Gone Absolutely Mental 💰 AI researchers are being recruited like Premier League strikers now. We're talking $250M packages. For a 24-year-old. I don't care how good your transformer architecture paper is, that number should make everyone uncomfortable. The maths works out to roughly "we'd rather overpay by 10x than let a competitor have you." Which, fine, that's how bidding wars work. But it tells you something about the state of things when the talent pool is so thin that a single researcher commands more than most companies are worth . Here's what bothers me though. These packages aren't salary. They're structured as equity, retention bonuses, and golden handcuffs. The researcher doesn't actually get $250M unless the company's valuation holds. And if we've learned anything from the last two decades of tech, valuations are vibes until they're not. Feels like: Paying someone a quarter billion to fix your plumbing while the rest of the house is on fire and nobody's rung the fire brigade. GPT-5: The Main Event (That Got Upstaged) 🎬 OpenAI shipped GPT-5 on Friday. Three variants: Pro, Mini, and Nano. Available to everyone in ChatGPT. Smartest and fastest, they say. And honestly? The reaction was a bit muted. Not because GPT-5 is bad. By all accounts it's proper good. But it landed in a week where everything dropped. Opus 4.1 on Wednesday. GPT-OSS on Wednesday. Gemini coding agent. Cursor CL

Steven Gonsalvez 2026-07-10 23:43 1 原文
AI 资讯 Dev.to

Vibe Coding Peak Hype: Windsurf Acquisition Chaos and the AI IDE Wars

Who Actually Bought Windsurf? 🌀 Three companies walk into a bar. They all claim they bought the same startup. Right. So here's the week in AI coding acquisitions, and I need you to stay with me because it gets properly daft. Monday : OpenAI's $3 billion bid for Windsurf collapses after Anthropic yanks their Claude API access. Brutal move, that. Like cutting off your rival's electricity mid-negotiation. Also Monday : Google DeepMind swoops in and hires Windsurf's CEO Varun Mohan, co-founder Douglas Chen, and the key researchers. Not an acquisition. A reverse acqui-hire. Price tag: $2.4 billion. For the people , mind you. Not the product. Tuesday : Cognition (the Devin lot) scoops up what's left: 250 engineers and $82 million in ARR. So in the space of 48 hours, Windsurf got rejected, stripped for parts, and sold off at a car boot sale. If you're a Windsurf user wondering what's happening to your IDE, the honest answer is: nobody has a clue, least of all Windsurf. Feels like: Watching a pub quiz team implode mid-game, with three other teams fighting over which one gets to adopt the remaining members. Steve Yegge Says What Everyone's Thinking The Pragmatic Engineer ran an interview with Steve Yegge this week, you know, the bloke who's been ranting about platforms since before most current SWEs had GitHub accounts. His take on vibe coding: it's deceptively hard . Not the coding itself (the AI handles that bit). The hard part is knowing whether what the AI produced is any good. He reckons there's an emerging "AI Fixer" role inside companies, someone whose entire job is reviewing and correcting AI-generated code. Which... yeah. That tracks. I've been saying this for months. Vibe coding is mint when you're prototyping or bodging together a script. But the second you need it to work in production, at scale, with actual users? You need someone who understands the code the machine spat out. And right now that someone is still a human. The hot take from the same week, "all mod

Steven Gonsalvez 2026-07-10 23:42 1 原文
AI 资讯 Dev.to

I Benchmarked 42 Compression Formats Spanning Four Decades. Here's What to Actually Use.

I run ezyZip , a browser-based archive tool, so "which format should I use?" is a question I field constantly. The honest answer is usually "it depends," which satisfies nobody. So I stopped hand-waving and measured it. We benchmarked 42 archive and compression formats, spanning four decades, from 1984's Unix compress through today's Zstandard, Brotli, and context-mixing paq8px. Everything ran against the same realistic 55 MB corpus, every archive was round-trip verified byte for byte, and the whole thing reproduces from a single command. Here's what came out of it, and what I'd actually reach for. The setup Most compression benchmarks measure raw codecs on standardized corpora like Silesia. That's the right call for algorithm research and the wrong call for answering "what should I zip my folder with?" I wanted end-user formats, real CLI tools, container overhead and all, on data that looks like an actual folder. So the corpus is deliberately mixed: about 11 MB of text, 15 MB of office documents, 16 MB of images, and 13 MB of video, all public domain so it can be committed and redistributed. That mix matters. Office documents ( .docx , .xlsx , .pptx ) are themselves ZIP containers, so they stress how a tool handles already-compressed data. The JPEG and H.264 media is near-incompressible and sets an honest lower bound. The plain text and uncompressed images are where formats actually separate. Two rules kept it fair and practical: Only two levels per tool: its default, and its one "maximum compression" dial. No method tuning, no dictionary sizes, no thread-count games. That's what a normal person can reach. Everything is round-trip verified. Each archive gets extracted, and every file is hashed with SHA-256 against the original manifest. Exit codes are not trusted. That last rule earned its keep immediately. The verification gotcha On the image category, a 1985-era ARC build produced an archive that its own extractor happily unpacked, while printing a CRC warning an

Andrew Dyster 2026-07-10 23:41 3 原文
开源项目 The Verge AI

Disney Plus is reportedly looking into a free streaming tier

Disney Plus is considering making some of its content free to watch, according to a report from Business Insider. A source tells the outlet that Adam Smith, Disney's chief product and technology officer, mentioned a free streaming tier during the company's town hall on Thursday. It's not clear which shows or movies the purported free […]

Emma Roth 2026-07-10 23:33 9 原文
开发者 HackerNews

Show HN: SubjectiveZero, an open-source agentic node editor for creative coding

Hey there, My name is Clem, I've been a solo indie dev for a couple years now, exploring frontier tech like XR and agentic workflows in the context of creative / interactive work. I've been building creation tools for a while and some common design challenge is to figure out the right level of abstraction for your tool. You can always make it super advanced and complex with low level concepts (shader composition, actual code etc.) but then you get something with a high complexity / learning curv

tasoeur 2026-07-10 23:23 2 原文