AI 资讯
Claude Opus 4.8 shipped today. Here is what the launch post does not say about why your agents will feel different tomorrow.
Claude Opus 4.8 shipped today. The benchmarks are a distraction — here is what actually changes about how your agents run tomorrow. Anthropic announced Claude Opus 4.8 at 16:00 UTC on June 3, 2026. The launch post leads with the usual benchmark deltas: SWE-bench Verified up 4.1 points, GPQA Diamond up 2.9, TAU-bench tool-use up 6.4. There is a chart. There is a marketing line about "the most capable agentic model we have ever shipped." If you stop reading there, you will miss the three things that will change how your production agents behave starting tomorrow. I have spent the morning re-running our internal agent harness against Opus 4.8 and reading the model card line by line. Two of the three changes are improvements. One of them is a silent regression that will bite anyone who pinned the model ID. Here is the full picture. What 4.8 actually changes The model card and release notes ship three changes that the launch blog post does not foreground: Cache-aware routing inside long agentic loops. The 4.7 router treated every tool-call cycle as a fresh planning step. 4.8 keeps an internal trace of which cache breakpoints were hit on the previous step and biases the next plan toward extending those traces. In agent harnesses that already use prompt caching aggressively (Claude Code, the Agent SDK with cacheControl: "ephemeral" on the system prompt), cache hit rates jumped from a measured ~46% on 4.7 to ~71% on 4.8 across a 30-step coding loop. The 200k context window now actually behaves at 200k. Anthropic published a needle-in-a-haystack chart in the model card going out to 200,000 tokens. The 4.7 chart got noticeably worse past ~140k tokens; the 4.8 chart is flat. This sounds like a benchmark thing. It is not. It changes the cost equation for "just stuff everything in context" patterns that 4.7 quietly punished by degrading accuracy. claude-opus-4-7 was not aliased. The launch shipped a new model ID — claude-opus-4-8 — and the previous ID is still callable. But if y
开发者
How many of us developers end up inadvertently running businesses? Starting out with hobby projects, seeing that there are operating costs, wanting to charge users or monetize and eventually having to register a company to funnel payments correctly?
开源项目
Meta will reportedly let employees take 30-minute breaks from its tracking program
Workers can pause the all-seeing eye when they need to "check something personal."
AI 资讯
Best Cryptocurrency APIs in 2026: Ultimate Guide for Developers and AI Agents
Crypto APIs are no longer just tools for fetching Bitcoin prices. In 2026, they are becoming the...
开发者
Is This How We'll Build Websites Soon? (webMCP Live Demo 🚀)
A few years ago, we started adapting our websites for mobile devices. Then we adapted them for...
产品设计
Cyera eyes $12B valuation at 80x ARR multiple despite operating losses
The cybersecurity company is nearing a $300 million round led by Evolution Equity Partners.
AI 资讯
God of War Laufey is coming to the PS5
Sony ended its big State of Play showcase with a major reveal: the next God of War. The new title is called God of War Laufey, and is once again developed by Sony's Santa Monica Studio. Currently, the game doesn't have a date, but it's coming to the PS5 whenever it does launch. This time, […]
AI 资讯
Remedy’s Control sequel launches in September
Control Resonant, the upcoming sequel to Remedy Entertainment's Control, will be released on September 24th, 2026, according to a trailer that premiered during Tuesday's PlayStation State of Play show. The trailer also gave a preview of the game's story, which stars Dylan Faden, who was imprisoned for much of the first game. He will be […]
AI 资讯
A Curated List of Articles About Modern Software Testing
Software testing is changing quickly. Teams are dealing with faster release cycles, more AI-assisted development, more complex browser behavior, and higher expectations around product quality. I collected a few practical articles that cover different parts of modern QA, test automation, developer workflows, and testing strategy. Recommended reads How to Test AI Agents for Tool Use, Memory, and Recovery Paths A practical framework for testing AI agents for tool use, memory retention, retries, and recovery paths, with concrete strategies for QA and engineering teams. How to Evaluate a Test Automation Tool for Shadow DOM, iframes, and Other Hard-to-Test UI Surfaces A practical buyer guide for evaluating test automation tools for shadow DOM testing, iframe testing, resilient selectors, and dynamic UI edge cases. How to Reproduce a Flaky Browser Test with Video, Logs, and Network Traces A practical workflow to reproduce a flaky browser test using video, logs, and network traces, then turn intermittent failures into repeatable bug reports. Endtest Review for Small QA Teams: Where Editable Test Flows Save the Most Time A practical Endtest review for small QA teams focused on editable test flows, maintainable test steps, and where no-code QA automation actually saves time. Editable Test Steps vs Generated Test Code: Which Holds Up Better After UI Changes? A practical comparison of editable test steps vs generated test code for UI change resilience, maintenance overhead, debugging, and team handoff, with guidance for QA and engineering leaders. Managed QA Services vs Staff Augmentation: What Changes in Ownership, Speed, and Cost A practical comparison of managed QA services vs staff augmentation, focusing on ownership, ramp time, communication overhead, cost, and maintenance risk. Automation Payback Period: How Long Does QA Test Automation Take to Break Even? Learn how to estimate the test automation payback period, model QA ROI, account for maintenance cost, and identify wh
AI 资讯
Stop Juggling 5 Tools , Python's uv Does It All (And It's Blazing Fast)
If you've been writing Python for more than a year, you know the ritual. A new project. A fresh terminal. And then: pyenv install 3.12.3 pyenv local 3.12.3 python -m venv .venv source .venv/bin/activate pip install pip --upgrade pip install -r requirements.txt Six commands before you've written a single line of code. And that's if nothing breaks. Enter uv a single binary that replaces pip , virtualenv , pip-tools , pyenv , and pipx . Written in Rust. 10–100x faster than pip. And honestly, one of the most pleasant tools I've used in the Python ecosystem in years. Let's dig into it. What Even Is uv ? uv is a Python package and project manager built by Astral , the same team behind ruff , the linter that everyone switched to and never looked back. The goal is simple: be the Cargo for Python . One tool, one lockfile, no friction. It's a standalone binary with zero Python dependencies, which means it works even before Python is installed. Installing uv # macOS / Linux curl -LsSf https://astral.sh/uv/install.sh | sh # Windows (PowerShell) powershell -ExecutionPolicy ByPass -c "irm https://astral.sh/uv/install.ps1 | iex" # Or via pip if you prefer pip install uv Verify: uv --version # uv 0.9.x The Speed Claim Is It Real? Yes. Embarrassingly so. Here's a timed comparison on Apple Silicon (Python 3.14): Operation pip / venv uv Create virtual env ~2 seconds 35 milliseconds Install FastAPI + deps (cold) ~12s ~1.2s Install with warm cache ~8s ~0.1s The warm cache case is where uv really shines it uses a global cache and hard-links packages into environments instead of copying them. If you've installed requests in any previous project, your next project gets it nearly instantly. Starting a New Project This is where uv feels like a completely different world: uv init my-api cd my-api That single command gives you: my-api/ ├── .git/ ├── .venv/ ← already created ├── .python-version ├── pyproject.toml ├── README.md └── main.py No separate python -m venv , no git init , no template c
AI 资讯
Hybrid RAG, No-Code AI Agent Memory, & Google Workspace CLI for Agents
Hybrid RAG, No-Code AI Agent Memory, & Google Workspace CLI for Agents Today's Highlights Today's top stories delve into advanced RAG techniques, focusing on hybrid retrieval strategies to overcome limitations of vector-only search, and explore practical solutions for equipping AI agents with long-term memory. Additionally, we highlight a new unified CLI that empowers AI agents to automate tasks across Google Workspace, streamlining workflow automation. Why Vector Search Alone Isn't Enough: Hybrid Retrieval for RAG (InfoQ) Source: https://www.infoq.com/articles/vector-search-hybrid-retrieval-rag/?utm_campaign=infoq_content&utm_source=infoq&utm_medium=feed&utm_term=global This article addresses a critical limitation in current RAG (Retrieval-Augmented Generation) frameworks: the over-reliance on pure vector search. While semantic vector search excels at understanding conceptual similarity, it often struggles with exact keyword matching or retrieving information from documents that lack strong semantic context but contain vital terms. The piece advocates for hybrid retrieval, a strategy that combines semantic (vector-based) search with lexical (keyword-based, e.g., BM25) search. This combination significantly enhances the recall and precision of retrieved documents, leading to more accurate and contextually relevant responses from large language models. For practitioners, understanding and implementing hybrid retrieval is essential for building robust, production-grade RAG systems capable of handling diverse queries and document types, thereby improving overall document processing and search augmentation performance. Comment: Anyone building serious RAG apps knows vector search has blind spots. Hybrid retrieval is a non-negotiable step for production, ensuring critical keywords aren't overlooked and improving overall response quality. Give your AI agent long-term memory with MCP (no code) (Dev.to Top) Source: https://dev.to/lrdeoliveira/give-your-ai-agent-long-term-me
AI 资讯
Presentation Slides for RubyConf Austria 2026 Talk "Frontend Ruby on Rails with Glimmer DSL for Web"
My talk “Frontend Ruby on Rails with Glimmer DSL for Web” went well at RubyConf Austria 2026 . Especially given that after the talk, Chad Fowler (the starter of RubyConf and famous book author of The Passionate Programmer , among other books) told me “good job”, and Obie Fernandez (a famous entrepreneur and book author of The Rails Way , among other books) told me he will try Glimmer DSL for Web because he doesn’t like React.js. Presentation Slides Direct Original Long Link: https://docs.google.com/presentation/d/e/2PACX-1vQ9oBnZpzK_eicVLGSqDmVzhsXsblONEKepnw5_xGHGXTM52JSjaS_ObYUJbx-zkb1M2ul9N2A2MnvU/pub?start=false&loop=false&delayms=60000&slide=id.g140fe579a5a_0_0 Glimmer DSL for Web GitHub: https://github.com/AndyObtiva/glimmer-dsl-web I ran a poll at the beginning of my talk, and everyone agreed that they love Ruby and that Ruby is superior to JavaScript, plus the majority indicated that they’d like to write less JavaScript and more Ruby during their Rails web development work. Several attendees told me my talk was great after the talk. Charles Nutter had me help him with his JRuby workshop afterwards by showcasing my other Glimmer project, Glimmer DSL for SWT , which runs on JRuby. In about 1 minute, I scaffolded a Hello World desktop app from scratch and then packaged it as a native executable on the Mac. Attendees were impressed. So, I’ve participated in presenting 2 events at this conference. I am very grateful for having such a great experience at RubyConf Austria 2026 overall, especially given that it uniquely included several classical/neoclassical/jazz concerts in between talks that entertained us and relaxed us. Chad Fowler concluded the conference with a beautiful Jazz piano and sax performance. Shout out to Hans Schnedlitz, Muhamed Isabegovic, and Zuzanna Kusznir (plus everyone who helped out) for organizing and hosting such a special Ruby conference!!! Original blog post version: https://andymaleh.blogspot.com/2026/06/rubyconf-austria-2026-frontend-r
AI 资讯
How to Integrate the OpenAI API into a Production Express App
Last year I helped a startup integrate the OpenAI API into their product. It was a chat feature — users could ask questions about their data and get natural language answers. The integration took about a day. Three days after launch, the founder messaged me: "Hey, something's wrong. Our AWS bill just showed an unexpected charge." It was $340. For three days. They had 60 users. The issue wasn't a bug — it was that production API usage looks nothing like a tutorial. The tutorial shows you openai.chat.completions.create() and returns a response. The tutorial doesn't show you what happens when users send 500-token messages, when they open 15 browser tabs each maintaining their own chat context, or when one user fires requests 30 times per minute because they think it's broken. This guide covers what the tutorials skip: rate limiting, token counting, cost guards, streaming, error handling with retries, and model selection. These aren't optional additions — they're what separates a demo from a production feature. Why Production Is Different Here's the gap between tutorial code and production code, stated plainly: Concern Tutorial Code Production Code Cost control Not mentioned Token counting, spending limits, model selection by task Rate limiting Not mentioned Per-user and per-IP limits to prevent abuse Error handling try/catch that logs to console Typed errors, retries with backoff, user-facing messages Response delivery Wait for full completion, return at once Streaming via SSE — response appears as it generates Context management Each request is independent Conversation history managed, truncated at token limit Secrets management API key hardcoded or in .env (no rotation) Rotation strategy, usage monitoring, per-feature keys Let's build a production-grade Express API that addresses all of this. We'll go layer by layer. The Architecture ┌─────────────────────────────────────────────────────────┐ │ CLIENT (Browser / Mobile) │ │ POST /api/chat { messages: [...] } │ │ GET
AI 资讯
Here’s seven bloody minutes of Wolverine on the PS5
At its big State of Play show on Tuesday, Sony shared new look Marvel's Wolverine, the next big title from Insomniac Games that's launching exclusively on PS5 on September 15th. Dressed in the iconic yellow Wolverine outfit, Logan slunk around and stabbed his way through a bunch of enemies, using his blades to take them […]
开源项目
Microsoft's Project Solara is an Android OS designed for agents instead of apps
Microsoft missed the boat on apps, so get ready for agents.
AI 资讯
PlayStation State of Play June 2026: All the news and trailers
While the majority of Summer Game Fest takes place over the weekend, Sony is getting a jump on things with its next State of Play showcase, which airs on June 2nd at 5PM ET. It’s coming at a pivotal time for the PlayStation brand: The PS5 keeps getting more expensive and Sony’s big bet on […]
AI 资讯
Microsoft Build 2026: The 7 biggest announcements
Microsoft just kicked off Build 2026 with a keynote from CEO Satya Nadella and other company leaders. As expected, it was filled with announcements, ranging from new Surface hardware to an always-on personal assistant and updates across Microsoft's in-house AI models. If you didn't watch the event live, you can catch up on all the […]
AI 资讯
Uber caps employee AI spending after blowing through budget in 4 months
Uber's cutback has occurred after the company had reportedly encouraged staff to use AI as much as possible.
AI 资讯
New Microsoft tool lets devs spin up AI behavior tests using text descriptions
Microsoft on Tuesday took the wraps off Adaptive Spec-driven Scoring for Evaluation and Regression Testing, an open source framework for spinning up AI evaluations.
AI 资讯
Your AI agents are authorized by vibes. Here's how to fix that.
The AI agent security community has been converging on a problem. A researcher recently ran an experiment — feeding a memory-retrieval framework 10 scenarios involving certificate operations: signing, issuing, revoking, delegating. The system retrieved the right memory 8 out of 10 times. It matched the external authorization gate 7 out of 10. The conclusion: metadata per item isn't enough. You need a separate authorization gate over the proposed operation. That conclusion is correct. But I want to show what that gate actually looks like when you build it — because the primitive already exists, and it's older than LLMs. The problem is authorization, not retrieval Most agent frameworks today invest in memory and observability. The agent can recall what it did before. You can see what tools it called. Logs, traces, dashboards. What they don't have is a cryptographically enforced answer to the question: was this agent authorized to do this, before it did it? Those are different problems. Retrieval tells you what the agent remembers about its permissions. Authorization tells you what it was actually granted — signed, tamper-proof, at dispatch time. An agent that retrieves "I have revocation permissions" from memory and then revokes a certificate it shouldn't touch is not an authorization failure at the retrieval layer. It's an authorization failure at the gate layer — because there was no gate. Certificates are that gate A certificate is a signed declaration of what an entity is authorized to do. Issued once, verifiable offline in ~1ms, revocable instantly. We've used them for TLS, for IoT devices, for code signing. The same primitive works for agents. The model is simple: Orchestrator issues a certificate at dispatch time The certificate carries the agent's identity and its exact scope in meta Every tool call goes through a gate that verifies the certificate offline On completion — or abort, or timeout — the orchestrator revokes it // Orchestrator — dispatch const { cer