今日已更新 144 条资讯 | 累计 37695 条内容
关于我们

标签:#Agents

找到 836 篇相关文章

AI 资讯

Why Your Agent's Search Results Look Right and Are Wrong: The Index Distribution Problem

Why Your Agent's Search Results Look Right and Are Wrong: The Index Distribution Problem You've built an agent. It has a search tool. You query it with something reasonable — a factual question, a comparison, a technical lookup — and it returns results. The results look right. The sources are real. The snippets are plausible. The agent synthesizes them into a confident answer. And the answer is wrong. Not obviously wrong. Not hallucinated-in-a-hallucinatory-way wrong. Structurally wrong — wrong in a way that passes every surface-level check because the error is baked into the retrieval layer before the model ever sees the context. This isn't a prompt engineering problem. It isn't a context window problem. It's a distribution problem , and it has a structural ceiling that no amount of better prompting will fix. The Index Is a Frozen Decision Here's the thing most agent builders don't internalize: a search index is not a neutral representation of knowledge. It's a frozen set of decisions about what matters and what doesn't. Every index — whether it's a BM25 inverted index, a dense vector store, or a commercial web search API — encodes a distribution shaped by past relevance judgments. Someone, at some point, decided which documents were "relevant" to which queries. That could be explicit (human raters labeling search results) or implicit (click logs, dwell time, link graphs). Either way, the index now encodes a probability distribution over what the system considers a good answer to a given query. That distribution is not semantic truth. It's past relevance consensus . Consider what happens when you embed a corpus and build a vector index. Your embedding model was trained on data that reflects certain assumptions about what concepts are close to each other. Your chunking strategy encodes assumptions about what granularity of information is useful. Your ranking model — whether it's cross-encoder reranking or a learned relevance model — was trained on labeled data that

2026-06-22 原文 →
AI 资讯

I opened my first PR to LiveKit's agents repo — here's the bug I found

I've been growing my open source portfolio one contribution at a time, and this week I landed on something genuinely interesting in livekit/agents (11k+ stars, the framework behind a ton of real-time voice AI agents). The bug If you're building a voice agent on a realtime model (OpenAI Realtime, xAI, Gemini Live), the model streams your transcription back in chunks. A single utterance can fire many user_input_transcribed events before it's final — token by token for OpenAI/xAI, or as one big interim blob for Gemini. If you want to react exactly once per utterance (say, show a "user is typing" indicator on your frontend via RPC), you need a stable key to correlate all those interim events together. That key already existed internally — InputTranscriptionCompleted carries an item_id . But when the framework re-emitted it upward as the public UserInputTranscribedEvent , the item_id was silently dropped — leaving consumers with no reliable way to dedupe across providers. The fix Small once you see it: add the field, forward it. class UserInputTranscribedEvent ( BaseModel ): transcript : str is_final : bool item_id : str | None = None # new ... def _on_input_audio_transcription_completed ( self , ev : llm . InputTranscriptionCompleted ) -> None : self . _session . _user_input_transcribed ( UserInputTranscribedEvent ( transcript = ev . transcript , is_final = ev . is_final , item_id = ev . item_id ) ) Two files, about 10 lines of real change. The actual work was tracing the event from the realtime model layer, through AgentActivity , up to AgentSession , to find exactly where the field got swallowed. The takeaway I didn't need to understand all of livekit-agents to land this — just one event's lifecycle, end to end. Small, well-scoped issues are the most achievable way into a big codebase, especially when someone's already mapped the territory in the issue itself. PR is up, CI green, waiting on review: github.com/livekit/agents/pull/6172

2026-06-22 原文 →
AI 资讯

Evaluating Kimi 2.5 vs Kimi 2.6: What happens to agent skills when the model gets smarter?

When a stronger model ships, there are two questions every skill author should want answered, and evals are the only honest way to answer either: Which skills just got absorbed? A model that now knows how to do X natively does not need a skill telling it to do X. Fewer skills to maintain, leaner context, lower cost. Which skills still matter? Behaviour-level guidance (conventions, preferences, project-specific workflows) is not something pretraining will fill in for you. Those skills should keep paying. Moonshot gave us early access to Kimi K2.6. We ran the Tessl agent skill evaluation harness on the same 21 skills and 100 paired scenarios against three solvers: Kimi K2.5, Kimi K2.6, and Claude Sonnet 4.5. A solver is the model whose output the grader scores; a paired scenario is the same task run twice per solver, once without the skill installed and once with it. These are early signals from one pre-release on one skill set. A deeper cross-model analysis with clean baselines across the board is in progress and will be its own piece. What does our setup look like? Scenarios and rubrics are held fixed across the two Moonshot runs. The only variable is the solver. Solver A: Kimi K2.5 Solver B: Kimi K2.6 Scenario generator: Claude Sonnet 4.5, up to 5 scenarios per skill, derived from each skill's SKILL.md Grader: Claude Sonnet 4.5, weighted-checklist rubric derived from the same SKILL.md Per skill × per solver: every scenario solved twice, baseline (no skill installed) and with-skill Per-skill n=5 is noisy; the aggregate over 100 scenarios is where the signal lives. Three findings: Kimi 2.6 is a better model than K2.5: Without skills, K2.6 sits ~2 pp (percentage points) above K2.5 in aggregate, with double-digit moves on specific skills. Kimi 2.6 holds its own against Sonnet 4.5. We picked Sonnet 4.5 as a competitive baseline, and found in this evaluation set that the K2.6 performed better both in the with/without skill scenario by around ~8 p.p . Skills remain a dura

2026-06-21 原文 →
AI 资讯

Goal In, DAG Out: How Open-Multi-Agent Turns a Goal into a Task DAG

You wrote the graph by hand. Then the requirements changed. Most TypeScript agent frameworks make you draw the graph yourself. You declare the nodes, wire the edges, decide what runs after what, where it branches, where it joins. It works, right up until the goal shifts and you are back in the graph editor re-wiring a pipeline you already built once. There is another way to model this: describe the goal, and let a coordinator build the graph for you. That is what runTeam() does in open-multi-agent. You hand it a team and a sentence. It hands back a result. In between, a coordinator agent decomposes the goal into a task DAG, assigns the tasks to your agents, runs the independent ones in parallel, and synthesizes the final answer. There are no edges to wire. This post is about what happens in that "in between," because the mechanism is the whole point. The one call import { OpenMultiAgent } from ' @open-multi-agent/core ' const orchestrator = new OpenMultiAgent ({ defaultModel : ' deepseek-v4-flash ' , defaultProvider : ' deepseek ' , }) const team = orchestrator . createTeam ( ' research ' , { name : ' research ' , agents : [ { name : ' researcher ' , model : ' deepseek-v4-flash ' , provider : ' deepseek ' , systemPrompt : ' You research topics and gather concrete facts. ' }, { name : ' writer ' , model : ' deepseek-v4-flash ' , provider : ' deepseek ' , systemPrompt : ' You turn research notes into clear prose. ' }, ], sharedMemory : true , }) const result = await orchestrator . runTeam ( team , ' Research the tradeoffs of TypeScript decorators, covering the stage-3 standard ' + ' versus the legacy experimental implementation, runtime and bundle-size cost, and ' + ' current framework support, then write a clear 500-word explainer for a team ' + ' deciding whether to adopt them. ' , ) console . log ( result . agentResults . get ( ' coordinator ' )?. output ) Three things to notice before we go under the hood: You never declared a task graph. You wrote the goal in pla

2026-06-21 原文 →
AI 资讯

Day 9 of building an AI agent that controls a phone. It works perfectly on my phone. But on a friend's phone, template matching failed. Icons rendered differently. The agent couldn't send a message. Now I'm exploring UI hierarchy inspection

Project Log #9: My AI Agent Works on My Phone. But What About Yours? Okeke Chukwudubem Okeke Chukwudubem Okeke Chukwudubem Follow Jun 20 Project Log #9: My AI Agent Works on My Phone. But What About Yours? # ai # webdev # programming # productivity 1 reaction Add Comment 3 min read

2026-06-21 原文 →
AI 资讯

60–95% fewer tokens in your agent loops, same answers. Meet Headroom.

AI coding agents are expensive — not because models cost too much per token, but because they send too many of them. An SRE debugging session with a raw agent: 65,694 tokens in. With Headroom in the middle: 5,118. Same bug found. Headroom is a new open-source context compression layer that intercepts everything your agent reads — tool outputs, log dumps, RAG chunks, files, conversation history — and compresses it before the LLM ever sees it. It's local, reversible, and available as a drop-in proxy, a library, or an MCP server. The numbers that matter Savings on real agent workloads: Code search (100 results): 17,765 → 1,408 tokens (92% reduction) SRE incident debugging: 65,694 → 5,118 tokens (92%) GitHub issue triage: 54,174 → 14,761 tokens (73%) Codebase exploration: 78,502 → 41,254 tokens (47%) Accuracy on standard benchmarks (GSM8K, TruthfulQA, SQuAD v2, BFCL) is preserved — some scores actually improve slightly, likely because the model sees cleaner signal. What's doing the compression Under the hood, Headroom routes content through a stack of specialised compressors: SmartCrusher — JSON, nested objects, arrays of dicts CodeCompressor — AST-aware for Python, JS, Go, Rust, Java, C++ Kompress-base — a custom HuggingFace model trained on agentic traces, for prose and mixed content CacheAligner — stabilises prompt prefixes so Anthropic/OpenAI KV caches actually hit It also does CCR (reversible compression) — originals are cached locally and the LLM can retrieve them on demand if it needs them. Nothing is destroyed. Why the proxy mode matters The most interesting deployment path: headroom proxy --port 8787 , then point your existing tool at localhost. Zero code changes. Works with any language. Or even simpler: headroom wrap claude wraps Claude Code, routes its traffic through Headroom automatically. One command, savings start immediately. Same for Codex, Cursor, Aider, Copilot CLI. "Library — compress(messages) in Python or TypeScript, inline in any app. Proxy — hea

2026-06-20 原文 →
AI 资讯

Give Your Codebase a Constitution

Architecture that lives only in people's heads doesn't survive agents. For most of my career, the real rules of a codebase weren't written down. People knew them. Senior engineers knew which layers could talk to which. They knew which dependencies were forbidden, which schemas were effectively frozen, and which shortcuts would create problems six months later. New engineers learned those rules the traditional way: break one, get caught in review, get the explanation, and eventually remember not to do it again. It wasn't perfect, but it mostly worked. What I didn't fully appreciate until I started working heavily with coding agents is how dependent that model is on tribal knowledge. Humans accumulate context over time. Agents don't. They don't remember the migration that went sideways three years ago. They weren't around when the team spent weeks untangling a dependency cycle. They don't know why a particular boundary exists. They only know what they can see. Which means if a rule isn't written down, from the agent's perspective, the rule doesn't exist. I've seen agents wire inner layers directly to outer layers. I've seen them introduce dependencies we intentionally avoided and extend contracts everyone on the team considered settled. The code often worked, which was the dangerous part. The problem wasn't correctness. The problem was architectural drift. That's when something clicked for me. Architecture can't remain folklore once agents start writing code. It has to become law. Not a convention. Not a suggestion. Not something a reviewer remembers at 6 PM on a Friday. A law. Written down, explicit, and enforceable. That's what I mean by a constitution. A Constitution Is Not Documentation The first mistake I made was treating the constitution like another documentation file. It isn't. Documentation explains how the system works today. A constitution defines what the system is allowed to become. Those sound similar, but they serve very different purposes. Package nam

2026-06-20 原文 →
AI 资讯

Hermes Agent Skills — Self-Evolving, Persona-Aware Skill Collection for Hermes Agent

Body: Hey everyone 👋 I've been building hermes-agent-skills — a production-grade skill collection for Hermes Agent that does three things no other skill pack does: 1. Self-Evolving Skills Skills aren't static YAML. The built-in EvolutionEngine tracks 5 health dimensions (usage frequency, success rate, user corrections, freshness, command validity), assigns a health score, and tells you which skills are rotting. Think of it as npm audit for your AI assistant's capabilities. 2. SOUL.md Persona Awareness Drop a SOUL.md in your Hermes config — naming conventions, comment density, architecture preferences, commit style — and every skill that touches code output adapts to it. hermes-skill soul generate bootstraps one in one command. The persona-aware-coding skill reads it at runtime so your agent writes code that actually looks like you wrote it. 3. CLI Toolchain hermes-skill create my-workflow # scaffold a standards-compliant SKILL.md hermes-skill validate skills/ # validate against the Agent Skills Standard hermes-skill list skills/ -f json # enumerate with health metadata hermes-skill soul generate # bootstrap a persona file What's in the box (v1.1.0): | Skill | Phase | Hermes-only Feature | |---|---|---| | requirement-analyzer | Define | Persistent memory across sessions | | spec-driven-dev | Spec | /skills chain forming workflows | | test-driven-dev | Build | delegate_task parallel test execution | | debugger-coordinator | Verify | browser + terminal + vision tri-tool | | code-quality-guardian | Review | patch auto-fix + /curator tracking | | cicd-orchestrator | Ship | cronjob scheduling + webhook triggers | | skill-curator | Evolve | Direct /curator integration | | persona-aware-coding | Identity | Native SOUL.md persona system | Why this is different: Most agent skill collections are portable but shallow — they can't use any platform's unique superpowers. These skills go deep on Hermes specifically: slash commands, delegate_task, persistent memory, vision+browser+t

2026-06-20 原文 →
AI 资讯

🚀 I Ran Claude Code on Every New Claude Model. Here's What Actually Ships.

Fable, Mythos, Opus 4.8, Sonnet 4.6, Haiku — Anthropic's 2026 lineup is no longer "one model you talk to." It's a fleet you route between. I spent a month inside Claude Code orchestrating all of them across real codebases. Here's which model to reach for, when, and the routing playbook that quietly doubled my throughput. Why I Went Down This Rabbit Hole (Again) Last time I wrote about Claude Skills and called Claude Code the killer host for them. Since then, two things happened that changed how I work day to day. First, the models got genuinely strange-good . In the span of a few months Anthropic shipped Sonnet 4.6, Opus 4.8, and then an entirely new tier above Opus — the Mythos class — released to the public as Claude Fable 5 . We went from "the AI suggested a decent diff" to Stripe reporting that Fable 5 ran a codebase-wide migration on a 50-million-line Ruby codebase in a single day — work that would've taken a team over two months by hand. Second, Claude Code stopped being a single-model tool. With a fleet of models at different price/speed/intelligence points, the highest-leverage skill in 2026 isn't prompting — it's routing . Knowing which model to put on which task is the difference between burning $200 of tokens on a typo fix and one-shotting a multi-service refactor. So I did the obvious thing: I wired all of them into Claude Code and ran them against real work for a month — bug fixes, migrations, greenfield features, test suites, the boring stuff and the scary stuff. This is what I learned. TL;DR The lineup is now a ladder : Haiku → Sonnet 4.6 → Opus 4.8 → Fable 5 → Mythos 5. Each rung trades cost for capability and patience for long-horizon autonomy. Sonnet 4.6 is your default. Frontier-ish coding at $3/$15 per million tokens with a 1M-token context window . Most of your work should live here. Opus 4.8 is the reliable senior. Better judgment, ~4× less likely to let its own code bugs slide, and it powers dynamic workflows — hundreds of parallel subagents i

2026-06-20 原文 →
AI 资讯

AI Model Failover Drills: Keep Agents Useful When Providers Break

A model fallback that only works in a diagram is not resilience. It is a TODO with better branding. If your product depends on AI agents, one slow provider, rate-limit spike, regional restriction, malformed response, or model behavior change can turn a useful workflow into a confusing user experience. The dangerous part is not always a clean outage. The dangerous part is a half-working fallback that silently changes schemas, drops tool state, skips citations, or gives users lower-confidence output without saying so. This guide shows how to run practical AI model failover drills before production traffic teaches you the lesson the hard way. The goal is not to make every model interchangeable. The goal is to keep the user workflow safe, honest, and recoverable when the primary model cannot do the job. Why model failover needs drills, not just retries Most teams start with a simple fallback chain: try the primary model, then a backup model, then show an error. That is better than nothing, but it misses the real problems in AI applications. Traditional APIs usually fail in obvious ways: timeout, 500, bad credentials, quota exceeded. AI systems can fail more subtly: The backup model returns valid JSON with different field meanings. A cheaper model ignores part of the tool policy. A provider accepts the request but streams tokens too slowly. A fallback model does not support the same function-calling format. A regional policy or access rule changes availability. The model completes the answer but loses citation discipline. The agent retries and burns the tenant budget. The final response looks polished but skipped the expensive verification step. Recent AI infrastructure conversations are pointing in the same direction: the system around the model now matters as much as the model. Agent benchmarks, provider reliability, AI cost pressure, and model routing are all active developer concerns. Search results also show many broad posts about LLM fallback strategy, but fewer pr

2026-06-20 原文 →
AI 资讯

your CI agent is reading more than your prompt

The dangerous thing about CI agents is not that they can write code. It is that they run in the place where we already concentrate trust. CI has repository access. CI has tokens. CI has build logs. CI can fetch dependencies, publish artifacts, comment on pull requests, open issues, deploy previews, and sometimes touch production systems. It is the automation layer we taught ourselves to trust because the alternative was humans doing the same boring steps by hand. Now we are putting agents inside it. That is useful. It is also exactly where the security model gets weird. Microsoft published a write-up this month about a Claude Code GitHub Action case where untrusted GitHub content and file-reading capability could combine badly. The short version is that an agent operating in a CI/CD context had enough ambient access to read more than the user probably intended, including process environment data that could expose workflow secrets. Anthropic mitigated the issue in Claude Code 2.1.128. The specific bug matters. The pattern matters more. CI/CD agents are not chatbots with a build badge. They are automated actors running in a high-trust environment while reading untrusted instructions from pull requests, issues, comments, commit messages, files, logs, and whatever else the workflow feeds them. That combination deserves more fear than it is getting. prompts are now part of the attack surface We are used to thinking about CI security in terms of code and configuration. Who can modify the workflow file? Which secrets are available to pull requests? Do forks get privileged tokens? Are dependencies pinned? Are artifacts trusted? Can a build script publish something? Does the workflow run on pull_request or pull_request_target ? Those questions still matter. But agents add another layer: text becomes operational input. The agent may read a pull request description. It may read a comment asking it to fix a test. It may read source files changed by an untrusted contributor. It

2026-06-20 原文 →
AI 资讯

Your AI Agent Isn't Broken. Your Company's Truth Is.

The AI agent had one job: pay approved vendor invoices, so the finance team could stop doing it by hand. On a Tuesday morning, it picked up invoice #4471 from a freight vendor Ksh48,000, stamped Approved in the company's ERP, cleanly matched to a valid purchase order. The agent checked the things it was told to check. They all passed. It paid the invoice. The invoice had already been paid. The previous Thursday. By a member of the finance team. Here is what the company's systems believed that morning and none of them was wrong. The ERP said: Approved. Unpaid. The reconciliation job that pulls in bank activity runs overnight, and last night it had failed silently. So the ERP's picture of the world was simply four days stale. The bank feed said: Paid. Last Thursday. It was right. Nobody had told the ERP. A Slack thread said: "hold everything to this vendor they double-billed us last quarter, I'm sorting it out with their AP team." Posted by the accounts-payable lead. Three days earlier. Resolved in her head, and nowhere else. The vendor's own email said: "Payment well received, thank you!" referring, of course, to Thursday's payment. The agent's inbox reader had seen it that morning, then set it aside, because email ranked below the ERP and the two disagreed. Every system was internally consistent. Every system was the authority on something . And there was no system anywhere not one that could answer the only question that actually mattered: has invoice #4471 been paid? A human clerk would almost certainly have caught it. Not because a clerk is smarter than the model they're not. Because a clerk would have felt the friction. They'd have half-remembered cutting the check. Or scrolled past the Slack message that morning and hesitated. Or simply had the reflex to ping someone before sending $48,000 out the door. Reconciling systems that quietly disagree is most of what operations people actually do all day so much of it that nobody files it under "work." It's just judgm

2026-06-20 原文 →
AI 资讯

Qwen3.6-27B + vLLM + Hermes on 24GB VRAM: May 2026 Recipe

If you want to reproduce my current local Hermes Agent + Qwen3.6-27B setup, this is the shape I would start from. Target One local coding agent. One 24GB GPU. Long context. Tools enabled. Thinking enabled. No child agents fighting the main request. The goal is not peak tok/s on a short prompt. The goal is: can the same agent session keep working after hours of tool calls without losing prefix locality, timing out during prefill, or getting wrecked by auxiliary requests? Model This setup is intentionally text-only. I am not serving the multimodal GGUF variant here. The working configuration uses groxaxo/Qwen3.6-27B-GPTQ-Pro-4bit through vLLM with --language-model-only . That choice matters. On a 24GB RTX 3090, the text-only GPTQ-Marlin path gave the best balance I found between long context, prefix caching, stable agent behavior and usable decode speed. Vision should be handled by a separate service/model if needed. vLLM The useful shape: CUDA_VISIBLE_DEVICES = 0 vllm serve groxaxo/Qwen3.6-27B-GPTQ-Pro-4Bit \ --served-model-name qwen3.6-27b-gptq-pro-4bit \ --dtype float16 \ --quantization gptq_marlin \ --tensor-parallel-size 1 \ --max-model-len 131072 \ --max-num-seqs 1 \ --kv-cache-dtype fp8_e5m2 \ --enable-prefix-caching \ --reasoning-parser qwen3 \ --enable-auto-tool-choice \ --tool-call-parser qwen3_coder \ --gpu-memory-utilization 0.95 \ --max-cudagraph-capture-size 32 \ --language-model-only I used a recent vLLM nightly, not an old stable image ( 0.20.1rc1.dev16+g7a1eb8ac2 ). The two flags people will want to argue about: --max-num-seqs 1 --max-model-len 131072 I use max_num_seqs=1 deliberately. With an agent, parallelism is not free. Title generation, context compression, retries, browser checks, tool calls and side jobs can all steal KV/cache locality from the main request. On one 24GB GPU I prefer one useful request over two requests sabotaging each other. 131k context is tight, but workable here. If your service OOMs, reduce context before adding MTP or enf

2026-06-20 原文 →
AI 资讯

How we built an internal data analytics agent

Qubot, our internal Copilot-powered analytics agent, allows any GitHub employee to ask questions about our data in plain language. Here's what we learned as we built it. The post How we built an internal data analytics agent appeared first on The GitHub Blog .

2026-06-20 原文 →
AI 资讯

Presentation: AI Agents to Make Sense of Data at OpenAI

OpenAI’s Bonnie Xu discusses Kepler, an internal AI data analyst agent built to query 600+ petabytes of data. She explains how they overcome context window limits using MCP, automated code crawling, and RAG. Xu also shares how their team leverages scoped semantic memory for self-learning and utilizes AST-based LLM grading to build a robust, regression-free evaluation pipeline. By Bonnie Xu

2026-06-19 原文 →
产品设计

Azure Functions Ships Serverless Agents Runtime at Build 2026

Azure Functions shipped a serverless agents runtime in public preview at Build 2026. Agents are defined in .agent.md markdown files with YAML triggers, MCP server access, 1,400+ connectors, and sandboxed execution. The Functions team confirmed to InfoQ that the runtime adds no cold start overhead and no billing premium beyond standard Flex Consumption. By Steef-Jan Wiggers

2026-06-19 原文 →
AI 资讯

Windows Platform Security and the Race to Secure AI Agents

In a new Windows Developer Blog post titled "Windows platform security for AI agents", Microsoft positions Windows as the trustworthy operating system for autonomous agents and introduces the Microsoft Execution Containers (MXC) SDK as the core of that strategy. The post argues that containment, identity and manageability must be built into the operating system. By Matt Saunders

2026-06-19 原文 →
AI 资讯

I let Claude Code run --dangerously-skip-permissions on my production DB. Here's what I changed.

Last Tuesday at 3am, a multi-agent loop hit 12K KV writes/minute and froze. The loop was a one-line counter bug. That part was fixable. What I found while tracing it was worse. I had --dangerously-skip-permissions enabled on a Claude Code session that was running D1 migrations. I thought it was pointing at staging. It wasn't — I'd misconfigured my env file reference, loading .env.production instead of .dev.vars . Claude didn't ask. The flag told it not to. The migration was ADD COLUMN , not DROP COLUMN , so no data loss. Survivable. But only barely. The thing I got wrong: I treated --dangerously-skip-permissions as "skip the annoying confirmation popups." It's actually "remove the only moment a human sees what command is about to run." Those are very different things. Turning the flag back off helps, but it doesn't constrain what Claude attempts — it just adds a prompt you'll click through anyway at 3am. What actually worked was adding a deny rule in .claude/settings.json : { "permissions" : { "allow" : [ "Bash(wrangler d1 execute * --local*)" ], "deny" : [ "Bash(wrangler d1 execute *)" ] } } The allow rule is more specific than the deny, so --local calls go through and everything else is blocked before execution. Over 2 weeks post-fix, Claude attempted zero production DB commands. Three deny events were logged — all from ambiguous prompts I wrote during fast context-switches, not from Claude going rogue. I ended up running three layers: the settings.json allowlist, a separate git worktree for migration work that physically contains only staging credentials, and a CLAUDE.md that instructs Claude to ask before anything touching production. The CLAUDE.md approach has a real caveat though — in long sessions the instructions lose weight as context grows. Anything critical needs to be restated in the prompt itself. I wrote up the full breakdown — including the worktree setup, the exact CLAUDE.md wording, and why MCP tool permissions behave inconsistently with the deny ru

2026-06-19 原文 →
AI 资讯

What is Generative AI? Understanding the Foundation of Modern AI Agents #2

Everyone is talking about AI Agents. But before you build an AI Agent, there is one concept you absolutely need to understand: Generative AI. Generative AI is the technology that transformed software from systems that simply follow rules into systems that can understand language, generate responses, reason through instructions, and assist users in a natural way. As part of my new course: Develop Your First AI Agent with Microsoft Foundry I published the first lesson where we explore the journey from traditional software to Generative AI and understand why modern AI Agents became possible. 🎥 Watch the video here: Why This Topic Matters Many developers jump directly into AI Agents, prompts, tools, and frameworks. However, without understanding the evolution of AI, it becomes difficult to understand: Why AI Agents exist Why Large Language Models are important Why prompts work Why tools are needed How modern AI systems actually operate In this lesson, we start from first principles and build the foundation required for the rest of the course. What You'll Learn Traditional Software For decades, software followed a simple pattern: Input → Rules → Output Developers explicitly defined every behavior. This worked well until humans started interacting with software using natural language. Why Rule-Based Systems Break Imagine building a dietician chatbot. Users might ask: What should I eat? Suggest a healthy breakfast. What foods contain protein? Can I eat oats daily? All of these questions are similar. Yet they are phrased differently. Supporting thousands of variations quickly becomes impossible with manually written rules. Predictive AI Machine Learning introduced a new approach. Instead of writing rules, we train models using data. Examples include: Spam Detection Fraud Detection Recommendation Systems Predictive AI can make decisions. But it still cannot create content. Prediction vs Creation A predictive model can answer: Fraud probability: 87% But can it explain why? Ca

2026-06-19 原文 →