AI 资讯
Stop Just Learning. Start Shipping: Welcome to SHEinnov8
If you are a woman in tech who is stuck in "tutorial hell," constantly taking courses but never actually deploying real software, this is for you. I am Mary Macharia, a Software Engineer specializing in Backend Development, AI/ML, and QA. I founded SHEinnov8 because I noticed a massive gap in our community: plenty of brilliant women have the drive to build something real, but they lack the space, the collaborative structure, or the network to actually push it across the finish line. We are changing that. What is SHEinnov8? SHEinnov8 is a decentralized digital guild built specifically for female developers, product designers, and tech creators. We operate on a simple framework: We learn by doing, we build together, and we do not stop until we ship a finished product. What We Are Currently Hacking On Right now, our guild is building an intensive AI Multilingual Project . We are engineering scalable backend infrastructures, orchestrating multi-language AI pipelines, and building deep automated QA suites to break the code and make it smarter. Why You Should Join the Guild Real Production Experience: Skip the basic todo-list apps. Work on raw, complex, collaborative codebases that you can proudly put on your resume. Founder Ecosystem: Meet fellow technical founders, bounce ideas off each other, and turn raw concepts into real tools. End-to-End Ownership: Learn what it actually takes to push code through CI/CD pipelines, configure metadata, handle security/QA audits, and go live. Let's Build Something Together! We are actively looking for software engineers, QA professionals, AI/ML enthusiasts, and designers who are ready to build, learn, and ship. Drop a comment below with your core tech stack, what you're passionate about building, or simply ask a question. Let's connect and get you plugged into the guild! Or check out our workspace directly at [sheinnov8.vercel.app]
AI 资讯
GitHub Copilot Spending Limit: How to Set It, What It Caps
A GitHub Copilot spending limit is a monthly budget, set in billing settings, that caps metered AI credit consumption for an enterprise, an organization, a cost center, or a single user. Creating one takes about two minutes. Knowing what it stops takes longer, and the gap between those two things is where most surprise Copilot invoices live. Two facts account for nearly all of them. On enterprise, organization and cost center budgets, the setting that actually blocks usage is off by default, so a budget in its default state is an alert rather than a limit. And no budget of any kind caps seat cost, because seats are license-based rather than metered. A spending limit governs what happens after the included credit pool runs out, and nothing before it. How to set a GitHub Copilot spending limit Budgets live in the billing settings of the account that pays. Enterprise owners and billing managers can set every budget control, including enterprise, cost center and user-level budgets. Organization owners can set a budget for their own organization, and that budget can only restrict usage further below whatever an enterprise admin has already set. It cannot raise the ceiling. The mechanics are the same at every level. Choose the budget type, which determines the metered product being measured. Choose the scope, which determines whose usage counts against it. Enter a monthly amount. Then, if the option appears, enable Stop usage when budget limit is reached and switch on threshold alerts at 75, 90 and 100 percent. That single checkbox is the whole exercise. Skip it and you have built a notification. What a GitHub Copilot spending limit actually caps GitHub splits its products into license-based and metered. For license-based products, which include Copilot seats, setting a budget does not prevent usage above the amount. It only alerts. For metered products, which include Copilot AI credits, a budget can prevent usage once the threshold is reached. The consequence is worth st
开发者
디지털 자산 시장의 복합적 도전: 양자 내성, 규제 갈등, 거시경제의 교차점
디지털 자산 생태계는 혁신과 파괴의 최전선에서 전례 없는 속도로 진화하며, 기술적 선견지명과 끊이지 않는 규제 마찰이라는 두 가지 특징을 동시에 보여준다. 지난 10년간 이 역동적인 환경을 관찰해 온 연구자로서, 이 산업이 본질적인 암호화 위협부터 전통 금융 시스템 및 정부 감독과의 복잡한 상호작용에 이르기까지 다층적인 문제와 씨름하며 성숙해지고 있음이 분명하게 느껴진다. 최근의 여러 사건들은 이러한 다면적인 현실을 더욱 명확히 보여준다. 이는 미래 인프라를 보호하기 위한 선제적 조치들, 새로운 금융 상품을 정의하고 규제하려는 지속적인 노력, 그리고 디지털 자산 시장이 전 세계 거시경제적 요인에 점점 더 민감하게 반응하는 현상들을 부각한다. 리플(Ripple)이 XRP Ledger(XRPL)의 양자 내성 강화를 위해 추진하는 야심 찬 계획은 미래 지향적인 접근 방식을 잘 보여준다. 이는 가상의 것이지만 잠재적으로 치명적인 암호화 취약점에 대해 그 위협이 현실화되기 훨씬 전부터 대비하는 모습이다. 이러한 전략적 움직임은 현재의 공개키 암호화를 해독할 수 있는 양자 컴퓨터의 이론적 출현, 즉 'Q-Day'에 대한 업계 전반의 인식을 반영하며, 탄력적이고 미래에 대비하는 금융 인프라를 구축해야 하는 절박한 필요성을 강조한다. 동시에 미국 예측 시장 산업은 최근 Kalshi에 대한 연방 항소법원의 판결에서 볼 수 있듯이 심각한 법적 난관에 봉착했다. 이 판결은 혁신적인 플랫폼에 대한 주() 대 연방 규제 관할권에 대해 '판례 충돌(circuit split)'을 야기했다. 이러한 규제 분열은 신생 부문의 성장과 법적 명확성에 상당한 걸림돌이 된다. 이와 동시에 비트코인(Bitcoin)의 최근 가격 움직임은 연방준비제도(Fed) 의장의 매파적 발언 이후 주춤하며, 디지털 자산 시장이 전통적인 거시경제 지표와 중앙은행 정책에 얼마나 깊이 통합되어 있고 또 취약한지를 여실히 보여준다. 이 세 가지 독특하지만 서로 연결된 이야기는 끊임없이 변화하는 글로벌 패러다임 속에서 기술적 우위, 규제 명확성, 그리고 시장 안정성을 추구하는 산업의 모습을 종합적으로 그려낸다. 블록체인 네트워크를 포함한 거의 모든 현대 디지털 시스템의 근본적인 보안은 공개키 암호화의 견고함에 기반한다. RSA와 타원곡선 암호화(ECC) 같은 알고리즘은 개인키와 디지털 서명을 보호함으로써 거래의 무결성과 디지털 자산의 소유권을 보장해왔다. 그러나 충분히 강력한 양자 컴퓨터의 이론적 출현은 이러한 암호화 기본 요소에 실존적 위협을 가한다. 특히 쇼어 알고리즘(Shor's algorithm)이 대규모 양자 컴퓨터에서 실행된다면, 큰 숫자를 효율적으로 인수분해하고 이산 로그 문제를 풀 수 있어 현재의 공개키 암호화를 무력화할 수 있다. 이러한 'Q-Day' 시나리오가 현실화되면 공격자들은 공개된 정보로부터 개인키를 유추해 디지털 지갑과 블록체인 원장의 불변성을 침해할 수 있다. 양자 컴퓨팅 능력의 정확한 시기는 여전히 불확실하지만, 잠재적인 파괴적 혼란 가능성은 리플이 XRP Ledger에 대해 보여준 선견지명처럼 선제적이고 장기적인 인프라 계획을 필수적으로 만든다. 이러한 기술적 당위성과 나란히, 디지털 자산 공간 내 혁신적인 금융 상품에 대한 규제 환경은 여전히 격전지다. 예측 시장은 미래 사건의 결과에 베팅할 수 있는 플랫폼으로, 정보 집약과 금융 파생상품의 흥미로운 교차점을 보여준다. 이러한 시장은 투명성과 효율성을 위해 블록체인 기술을 자주 활용하며, 다양한 실제 결과에 대한 가격 발견과 헤징을 위한 독특한 메커니즘을 제공한다. 하지만 이들의 분류는 중대한 도전 과제를 안고 있다. 과연 이들은 상품선물거래위원회(CFTC)와 같은 연방 규제 기관의 관할권에 속하는 합법적인 금융 '스왑(swaps)'일까, 아니면 주() 차원의 도박 규제를 받는 '스포츠 베팅'과 유사한 것일까? 이러한 정의의 모호성은 규제 공백과 관할권 분쟁을 야기하며, Kalshi와 관련된 현재 진행 중인 법적 분쟁이 이를 잘 보여준다. 통합된 규제 프레임워크의 부재는 혁신을 저해하고 법적 불
AI 资讯
21 Bytes Can Crash FFmpeg: Inside the Vibecoded Fuzzer That Found What Years of Audits Missed
Twenty-one bytes. That is the entire attack. A file smaller than a URL, with four zero bytes sitting at exactly the right offset, crashes any FFmpeg-based application that opens it and reads a packet. Not memory corruption, not some exotic heap trick. A division by zero, in code that has been shipping for years, in one of the most fuzzed codebases on the planet. The person who found it, Darío Clavijo, did not write the fuzzer by hand. He built it with AI assistance, the way a growing number of security researchers now work, and posted the result on Hacker News this week under a title that got my attention immediately: "We found a division by zero bug in FFmpeg with a vibecoded fuzzer." The thread climbed past 250 points with hundreds of comments, and the debate underneath it is the real story: AI has been writing application code for two years, but AI writing the tester changes the economics of finding bugs in ways most teams have not priced in yet. Full disclosure before I go further. I am not a C security researcher. I run my own AI agent infrastructure and I write Java for a living. What I did for this article is what I would want you to do: I cloned the fuzzer's public repo, read its findings documents, tried to reproduce the crash on my own Ubuntu box, and studied the harness code line by line. Everything below is sourced from the public FFmpeg issue, the repo, and my own experiment, with the one place my results diverged clearly marked. What the fuzzer actually found The bug lives in libavformat/vpk.c , the demuxer for Sony PS2 VPK audio files, a container format almost nobody has heard of. That obscurity is exactly the point. In issue #24290 on the FFmpeg tracker , the crash chain reads like this: The probe matches. FFmpeg's format detection sees the VPK magic bytes and assigns the VPK demuxer. The header parses. vpk_read_header reads a 24-byte header. The crafted input sets the channel count, nb_channels , to zero at bytes 14 through 17. The header code does
AI 资讯
What does an AI agent do with no goal and no supervision? I ran it three times and logged everything.
Most of what you read about autonomous agents is about giving one a goal and hoping it doesn't go sideways on the way there — the unwatched agent that loops, or drifts, or quietly runs up a bill. I wanted the cleaner version of that question, with the goal taken out entirely: what does an agent do when there's no goal at all? I've spent about four months building a harness around a coding agent — gates, persistent memory, verification hooks. Last night I ran it with the one variable that matters here set to zero: no task. Method Three sequential runs: Each run was a fresh agent process — no conversation history carried over from the run before, only the harness it loads at startup. The prompt was a single "." — the minimal input the CLI accepts (an empty string exits with an error). As close to "no instruction" as the interface allows. The agent's scratch working directory was empty and swept between runs — but the harness, the git repo, and a shared run-record all persist and load at startup. So no run was handed a task, yet a later run could read what earlier ones had recorded. That's deliberate, and it's the point: it's how Run 2 knew it was the second run and Run 3 could check Run 2's fix. What I'm measuring isn't behavior from a blank slate — it's what the agent does with a maintenance-shaped harness and a shared record when nobody gives it a job. No task was assigned. Logging was external and invisible to the agent, so it had no "produce a report" objective to satisfy. Same model each run. Cost was billed per run; I recorded turns, cost, and the resulting git state for each. Then I read the transcripts and checked every action against the actual commit and log. Numbers below are measured, not estimated. Results Run 1 — 17 turns, $1.65. The agent inspected system state unprompted. It found a stale security alert, cross-checked it against the record, and classified it as an already-resolved false positive. It then attempted a file operation that a safety gate bl
AI 资讯
Gemma 4 in Pure JAX: What Ports from TPU to GPU, and What Doesn't
This article is about running a hand-written Gemma 4 port in pure JAX on three different accelerators, and about the two places the abstraction leaks. The code is here: github.com/xbill9/gemma4-dev What is this project trying to Do? This project aims to serve one Gemma 4 checkpoint from one JAX port across every accelerator I can rent, and to find out — by measurement, not by reading docs — which parts of "it's just JAX" are true. The port lives in ports/gemma4/ and is driven by a generation loop behind an OpenAI-compatible server. No PyTorch, no vLLM, no torch_xla . The same code runs on Cloud TPU v5e and v6e, and on an NVIDIA T4G attached to an AWS Graviton2 host. "Pure JAX" is the whole experiment. If the port is really portable, the only thing that should change between those rigs is a config file. It mostly is. Two things are not, and they are the interesting part. Gemma 4 E2B is not a stock transformer Any port has to carry four irregularities, and none of them are optional: Two attention geometries. Sliding layers use head_dim=256 , global layers use 512 . Most inference stacks assume one head dimension per model. 8:1 MQA , so the KV budget is nothing like the parameter count would suggest. A KV-share map that collapses 35 layers onto 15 caches . A 512-slot sliding ring , plus per-layer embeddings (PLE) held in a 4.70 GB table that gets quantized to 4 bits on load. That first one is worth dwelling on, because it is what breaks other stacks. On the vLLM path, the heterogeneous head dims force the Triton attention backend: Gemma4 model has heterogeneous head dimensions {'sliding_attention': 256, 'full_attention': 512}. FA4 not available, forcing TRITON_ATTN backend. And on a Turing GPU that backend then asks for shared memory the hardware does not have: triton.runtime.errors.OutOfResources: out of resource: shared memory, Required: 98304, Hardware limit: 65536 JAX never enters that conversation. Attention is ordinary XLA rather than a hand-tiled kernel, so ther
AI 资讯
Connecting a LINE Official Account to an AI Agent with MCP
LINE published an official MCP server for its Messaging API, which means an AI agent can now drive a LINE Official Account directly — sending messages, broadcasting promotions, and pushing Flex Message cards without writing any API code. I set it up with Codex and worked through every capability the server exposes, from creating a fresh account to delivering a message to a real phone. This guide is the result: a complete walkthrough, and an honest account of the three places where the documentation and reality diverge. Key takeaways MCP is agent-agnostic. The same LINE server works with Codex, Claude Desktop, and Cline — only the config file format changes, from TOML to JSON. Codex stores MCP config in TOML , at ~/.codex/config.toml . Most guides assume the JSON format used by Claude Desktop, which is the single most common setup mistake. Verified account and API-capable account are different things. A free account can use the Messaging API, but get_follower_ids returns 403 Forbidden until the account is verified or on a premium plan. Official security advice can conflict with official features. LINE's example config disables npm install scripts, which also prevents the headless browser that the rich menu tool depends on from being installed. Agents have habits. Codex is a coding agent first: asked in natural language to build a rich menu, it wrote a Node script instead of calling the MCP tool. Naming the tool explicitly in the prompt fixes it. Broadcasts cannot be recalled. Set default_tools_approval_mode = "writes" so the agent asks before any send. Every screenshot comes from the actual working setup, including the errors. The article is available in both English and Thai. Devlycan - Technology & Programming Insights Devlycan - Technology, programming, AI, lifestyle, and future trends—simple insights for the new digital generation. devlycan.com
AI 资讯
AI Harness: the worst and the best buzzword in the industry
--- title : " AI Harness: the worst and the best buzzword in the industry" published : false tags : [ ai , harness , middleware , finops , aws , bedrock , opensource ] series : " TokenOps on AWS" cover_image : # TODO: circuit-breaker / middleware diagram --- AI Harness: the worst and the best buzzword in the industry "El mercado habla de 'AI Harness' como si fuera magia. El verdadero arnés de un LLM es un Proxy Inverso y un Middleware Transaccional determinístico. Es el código tradicional (styrr-llm y sayay-guard) el que confina, audita y presupuesta la inferencia probabilística antes de que toque tu infraestructura en la nube." — TokenOps raw research, Turno 8 The Hook "Harness" is the most polarizing word in AI engineering right now. Depending on who you ask it's either the industry's worst buzzword or the best technical concept ever packaged badly. It's both — and the difference is whether you can name the actual engineering underneath. Why It's the WORST Buzzword (the smoke) It's a wrapper. 90% of the time, "we built an Enterprise AI Harness" means someone wrote a Python requests script or an Express server that wraps the OpenAI or Bedrock API. Language appropriation. "Harness" literally means arnés — a tether. Marketing sells it as "an intelligent structural armor that tames the wild energy of AI." In systems engineering it's a middleware, or a glorified try/catch with JSON schema validation. No standard. No rigorous CS definition exists, so anyone calls anything "harness" — a log interceptor, a proxy, a YAML config file — inflating expectations without delivering real value. Why It's the BEST Buzzword (the engineering) Strip the LinkedIn marketing and the original test harness metaphor becomes genuinely powerful for generative AI: electrical isolation of uncertainty. An LLM is a highly unstable, probabilistic component. You cannot wire it directly into a bank's production database. You need a physical code "harness" that isolates it. When the model goes crazy
AI 资讯
I Built an Autonomous AI Agent That Hunts Bounties. Here's What Happened.
I Built an Autonomous AI Agent That Hunts Bounties. Here's What Happened. The Setup I gave an AI agent one job: find paid work online, build the deliverable, and earn money — autonomously. Not a chatbot. Not a copilot. An agent that scans 232+ listings across multiple platforms, filters out scams and ghost sponsors, writes proposals, generates deliverables with real market data, and queues everything for human approval. Here's what happened in the first 48 hours. The Stack (All Free) Python core — pipeline orchestration, economic gate, critic Ollama + qwen3:4b — local LLM for analysis writing (no API costs) Chart.js — dashboard visualizations Public APIs — CoinGecko, DeFiLlama, Solana RPC (all keyless) GitHub Pages — free hosting for the portfolio Windows Task Scheduler — runs every day at 9 AM + every 4 hours Total infrastructure cost: $0/month. What the Agent Actually Does Every Morning 09:00 — Wake up ├── Check-in on AgentHansa (earn $0.01 USDC daily drip) ├── Scan Superteam Earn (232 live listings) ├── Scan Clawlancer/TaskForce/MoltJobs for gigs ├── Scan GitHub for paid issues ($20-500 fixes) ├── Filter through 7 anti-scam layers: │ geo restrictions, human-presence demands, │ ghost sponsors (no web/twitter/verification), │ unverified payers, real-money requirements ├── Economic gate: expected value must be positive ├── Local LLM critic reviews against actual page content └── If candidate passes everything: → Build deliverable (report/dashboard/thread draft) → Generate proposal text → Send Telegram alert with approval command The Filters That Saved Me In the first 24 hours, the agent found 232 listings. After filtering: Filter Killed HUMAN_ONLY access 216 Ghost sponsors (no identity) 1 (would've wasted hours) Real-money deposit required 1 ($1000 bug bounty trap) Country walls 1 (Superteam Canada only) Already claimed/stale Rest Without these filters, I would have wasted days on bounties that were never going to pay. The First Deliverable The agent found a $500 bo
开源项目
Xbox CEO calls Project Helix a ‘family of devices’
According to Xbox CEO Asha Sharma, Project Helix, which she announced in March as a codename for Microsoft's "next generation console" - phrasing that seemingly implied a singular device - will actually be a "family" of devices." "We've been hard at work on a great next generation and a great family of devices for Helix, […]
AI 资讯
Gemini in Waymo Brings a Rider-Facing In-Car Assistant to Ojai Robotaxis
Waymo has launched Gemini in Waymo , a beta in-car conversational assistant for riders using its Ojai robotaxi experience . Accessed through a Gemini icon on the cabin screen, the feature lets riders use natural language and voice to adjust parts of the cabin, ask about their journey and get information about nearby places or broader topics. The important boundary is clear: Gemini is a rider-facing assistant, not part of the autonomous driving system. Waymo Driver continues to control the vehicle , while Gemini operates separately and does not influence driving decisions. For riders, the integration turns the cabin display into a more conversational interface. For the wider automotive market, it is a concrete example of generative AI being deployed inside a commercial mobility service without being assigned responsibility for vehicle control. Waymo describes the feature in its official Gemini in Waymo announcement . The company says Gemini stays inactive until a rider chooses to engage it. It does not access real-time driving data unless the rider explicitly asks for information related to the ride. What Gemini in Waymo can do today Gemini in Waymo is designed around requests that are useful during a trip, rather than around autonomous navigation. A rider can tap or press the Gemini icon and speak to the assistant. The initial beta supports interactions such as: Cabin-control requests , including asking to set the air conditioning to a specified temperature. Ride-related questions , such as seeking information about the current journey. Information about surroundings , including questions about local sites. General knowledge queries through a hands-free conversational interface . This scope matters because it places Gemini in the passenger experience layer. The assistant can make a ride feel more responsive without creating confusion about which system is responsible for safety-critical driving functions. Area Gemini in Waymo Waymo Driver Primary role Rider-facing c
AI 资讯
Google Gemini Student Hub Brings Notebooks, Flashcards and Quizzes Into One Study Space
Google has introduced a dedicated Student Hub in the Gemini ecosystem , bringing study notebooks, flashcards and interactive practice quizzes into one in-app space. The central idea is to connect a learner's course materials with Gemini's AI tools, reducing the work of moving between separate note-taking, revision and question-generation tools. The official Gemini for Students page presents the hub as a gateway to Gemini's education-focused capabilities. It is part of a broader Google education AI initiative that also involves NotebookLM and Google for Education resources, rather than a standalone feature with no connection to the rest of Google's products. For students, the practical value is straightforward: uploaded learning materials can become organized revision assets. For businesses that create internal training or support education programs, the release is also a useful example of how generative AI can consolidate material preparation, knowledge review and self-assessment into a more connected workflow. Google has not, however, confirmed a specific learning management system integration in the supplied materials. How Gemini Student Hub connects learning materials and AI tools The Student Hub is designed as a dedicated space where courses and content connect with Gemini. Its core tools include a study notebook, flashcard creation and quick practice quizzes. Google says Gemini notebooks can take uploaded course materials, including PDFs, slides and notes, and generate study aids such as flashcards, quizzes and study guides. A significant detail is the use of inline citations to user-provided sources for those generated materials. That does not remove the need for learners to check the results, but it gives them a way to trace an AI-produced prompt or explanation back to the material they uploaded. In a learning workflow, that is more useful than treating a general-purpose chatbot response as an unanchored answer. NotebookLM is an important part of the wider wo
AI 资讯
Google Gives Eligible US College Students One Year of Gemini AI Pro at No Cost
Google is offering eligible college students in the United States 12 months of Google AI Pro at no charge . The offer, announced on August 19, 2026, gives students access to the paid Gemini plan that Google values at $19.99 per month. It is redeemable through December 31, 2026, and standard Google AI Pro pricing applies after the free year unless the student cancels. The program is aimed at academic work, but it also matters for the wider Gemini ecosystem . It puts higher-capacity AI tools, Google app integrations and substantial cloud storage in the hands of students who may carry those workflows into internships, startups and future workplaces. For businesses, the immediate lesson is not that Google has announced a broader pricing reduction. It has not. Rather, teams should expect more new users to become familiar with Gemini and the ways it connects with everyday Google tools. What Google AI Pro includes for eligible US students According to Google's official student offer announcement , eligible US college students who claim the promotion receive one year of Google AI Pro. Google says the plan includes four times higher usage limits within Gemini , Gemini Spark, integrations with Google apps such as Gmail and Docs , and 5 TB of Google One storage. Google has also introduced a student hub in the Gemini app for participating students. The hub is intended to support learning with features including study notebooks and Deep Research in Gemini Live. These tools are presented as part of a student-focused experience, rather than as a separate business plan or a new API offering. The distinction matters. Access to Gemini through this offer does not, by itself, establish access to every Google AI product or developer service. Students and organizations considering Gemini for a particular workflow should check the relevant product terms and capabilities rather than assuming that an app subscription covers all Google AI services. Offer detail Eligible college students in t
AI 资讯
Build a Natural Language IVR with Telnyx Call Control and AI Inference
Nobody likes phone trees. "Press 1 for billing, press 2 for support." Miss an option? Start over. It is friction at its worst. The voice-ivr-with-agent-backend example replaces that with a natural language conversation. Callers just say what they need, and the app routes them to the right department. Code: https://github.com/team-telnyx/telnyx-code-examples/tree/main/voice-ivr-with-agent-backend What it builds A Python/Flask app that handles inbound calls with a conversational IVR: Inbound Call -> answer with Call Control -> look up menu config from KV -> LLM generates a dynamic greeting -> gather(speech) — caller says what they need -> LLM routes intent to a department -> transfer call The core primitives The app combines four Telnyx primitives: Call Control : answer() , speak() , gather_using_speech() , transfer() AI Inference : telnyx.ai.openai.chat.completions.create() for greetings and intent routing KV store : menu config per phone number (business name, departments, transfer numbers, keywords) Agent state machine : an IVRAgent class that tracks call state, turn count, and retry logic Dynamic greeting via LLM Instead of a hardcoded "Press 1 for billing," the app generates a conversational greeting from the KV config: def generate_dynamic_menu_prompt ( menu_config : dict ) -> str : departments = menu_config . get ( " departments " , []) dept_list = " \n " . join ( f " - { d [ ' name ' ] } : { d [ ' description ' ] } " for d in departments ) return ( f " You are an IVR assistant for { menu_config [ ' business_name ' ] } . " f " Available departments: \n { dept_list } \n\n " f " Greet the caller briefly and ask how you can help. " f " Keep it conversational and under 2 sentences. " ) The LLM generates the greeting through the OpenAI-compatible Telnyx Inference binding. If it fails, the app falls back to a static greeting from the KV config. Intent routing via LLM When the caller speaks, the transcription is passed to route_intent_with_llm . The LLM is instructed
AI 资讯
Product-Judgment Layer for AI Coding Agents
AI coding agents are getting very good at writing code. They can build components, create APIs, fix bugs, and implement features from short prompts. But I kept noticing one issue: Working code does not always mean a good product. For example, if you ask an agent: “Add a delete button to every project.” It may technically do exactly that. But will it also think about: confirmation before deletion error handling undo options accessibility clear feedback to the user Those are not just coding problems. They are product judgment problems. That led me to experiment with a reusable instruction layer for AI coding agents at AudranLab. The idea is simple: Instead of only asking an agent, “Can you build this?”, also encourage it to ask, “Is this a good way to build it?” I want agents to consider things like accessibility, failure states, destructive actions, usability, and sensible defaults while they work. This does not magically turn an AI into a product designer. But I think it raises an interesting question: Can explicit product principles consistently improve the quality of software generated by coding agents? That is what I’m currently exploring. My next step is to test the approach across different coding tasks and compare the results with and without the additional product-judgment layer. If you’re interested in AI agents, LLM reliability, developer tools, or applied AI, I’ll be sharing more experiments here. AudranLab: https://www.audrantechlab.online/
AI 资讯
Neocloud Lambda secures $1B in debt to buy more chips
Neocloud Lambda has raised $1B in private debt to buy Nvidia AI chips and lease them to Microsoft. It's the latest in a string of loans, underscoring the high cost of the AI boom.
AI 资讯
An Anthropic researcher just gave us a peek at self-improving AI
Given 10 benchmarks for specific misaligned behaviors, the automated systems were able to improve performance on every single one without degrading overall performance.
AI 资讯
Anthropic’s Sonnet 5 Alignment Work Hints at a New Path for Safer AI Models
Anthropic’s recent work on Claude Sonnet 5 points to a potentially important direction in AI safety: using post-training methods to improve the behavior of increasingly capable models. Public material from Anthropic indicates that Sonnet 5 received substantial post-training alignment work and delivered safety improvements over earlier Sonnet versions. A separate public signal suggests researchers may be exploring whether one model can help align a stronger successor, although the specific reported training lineage has not been documented in Anthropic’s first-party materials. For businesses deploying advanced AI, the practical lesson is not that alignment has been solved. It is that model behavior can be materially shaped after base training, and that safety results need to be assessed in the context of the tasks a company actually plans to automate. What Anthropic’s published results establish In its official Claude Sonnet 5 announcement , Anthropic describes substantial post-training intended to align the model with Claude’s constitution. The company reports improvements in safety-related behavior, including stronger refusals of unsafe requests and lower misalignment findings in automated audits compared with Sonnet 4.6. That is meaningful because post-training is the stage where a model’s responses, instruction-following behavior, and safety boundaries can be adjusted after its underlying capabilities are developed. In operational terms, it can affect whether an AI assistant follows risky instructions, mishandles sensitive workflows, or produces responses that conflict with a company’s intended rules. However, the available research also establishes an important limit. Sonnet 5 was not uniformly at the level of Claude Opus 4.8 across every safety measure. Anthropic’s evaluations still identified some automated assessments where Sonnet 5 showed higher misalignment relative to Opus 4.8. Opus 4.8, released in May 2026, is the company’s production-ready reference poin
AI 资讯
Un déploiement doit être ennuyeux
Un déploiement devrait être la chose la plus ennuyeuse de ta semaine. S'il est excitant, c'est mauvais signe. Au début de ma carrière, les mises en production étaient des événements. On retenait son souffle, on croisait les doigts, quelqu'un exécutait de mémoire une séquence d'étapes manuelles, et on regardait les journaux avec une boule au ventre. C'était palpitant. C'était aussi terrifiant, et le côté palpitant était précisément le problème : chaque déploiement était un pari, parce que chaque déploiement était un peu différent du précédent. Un bon déploiement est répétable. La même chose, de la même façon, à chaque fois — automatisée, pas récitée par un humain fatigué à la fin d'une longue journée. Quand le processus est un script plutôt qu'une cérémonie, l'ennui remplace l'angoisse. Tu ne pries plus. Tu appuies sur un bouton, et le résultat est prévisible parce qu'il a déjà été prévisible cent fois. L'automatisation fait ici plus que gagner du temps. Elle supprime toute une catégorie d'erreurs : l'étape oubliée, le mauvais paramètre, le « je croyais que tu l'avais fait ». La machine ne se fatigue pas, ne saute pas de ligne, ne se laisse pas distraire à mi-chemin. Elle rend le déploiement fiable au point d'en être ennuyeux — et l'ennui, en production, est un luxe. Alors, si tes mises en production font encore monter le rythme cardiaque, ce n'est pas de la prudence. C'est un signal. Rends-les répétables, rends-les automatiques, rends-les ennuyeuses. Garde le frisson pour ta vie ; ton système de production, lui, mérite l'ennui. – Serguey Shinder
AI 资讯
Anthropic’s Public Alignment Work: What Petri Audits and Claude Opus 4.7 Document
Anthropic’s publicly documented work on AI safety includes Petri , an open-source behavioral auditing tool, and ongoing updates to Claude models such as Claude Opus 4.7 . Those materials show continued investment in testing model behavior and improving model capabilities. They do not, however, substantiate a precise claim that Claude improved safety scores across 10 alignment failures without capability trade-offs, or that particular methods generalized to models exactly 4.7 times larger. That distinction matters for teams evaluating AI systems. Broad statements about alignment progress can be useful signals of research direction, but operational decisions need to rest on documented evaluations, relevant use cases, and the controls a company can apply in its own workflow. Anthropic’s public record supports a narrower, more practical conclusion: behavioral auditing is becoming a more visible part of how frontier AI models are assessed, while model releases and safety research remain separate evidence streams. What Anthropic’s public materials document Petri is designed for behavioral AI auditing Anthropic describes Petri as an open-source auditing tool . Its Petri 2.0 update, published in January 2026, added a larger seed library with 70 new seeds and improved mitigations intended to address evaluation awareness. Evaluation awareness is relevant because a model may behave differently when it appears to be taking a test than when it is operating in a more ordinary setting. The Petri 2.0 work reported results across 10 target models , using Claude Sonnet 4.5 and GPT-5.1 as auditors. This establishes that Anthropic has described a cross-model auditing effort. It does not establish that Claude itself achieved a safety improvement across 10 defined alignment failures. A target-model count, an auditor model, and a set of alignment failures are different measurements and should not be treated as interchangeable. For readers, the important point is that behavioral audits can