今日已更新 408 条资讯 | 累计 28312 条内容
关于我们

今日精选

HOT

最新资讯

共 28312 篇
第 695/1416 页
AI 资讯 Dev.to

Why Organizations Need an AI Gateway

An AI gateway is the control point between your applications and the LLMs they call. It’s where cost, security, reliability, and governance get managed across every model and provider at once. Skip it, and AI sprawl quietly turns into runaway spend, security gaps, and outages you didn’t see coming. Here’s why a gateway has become core infrastructure. Almost nobody adopts AI in a tidy, planned way. One team ships a support chatbot on OpenAI. Another prototypes on Anthropic. A third fine-tunes an open model on its own GPUs because the latency was better. A year later you’ve got dozens of applications, several providers, API keys scattered across repos, and no single answer to a simple question: what are we spending, and what data are we sending where? That’s the gap an AI gateway fills. It sits between your applications and the models, and it turns fragmented, ungoverned access into something you can actually manage. The reason organizations end up needing one is straightforward — production AI creates problems that application code was never designed to solve. Let’s walk through them. The problems an AI gateway solves Cost that’s invisible until the invoice arrives LLM spend is uniquely easy to blow up. A retry bug, an agent stuck in a loop, an unbounded batch job — any of these can multiply tokens overnight. And when every team holds its own provider key, finance gets one large number with no story behind it. A gateway changes that. It enforces budgets and rate limits per user, team, and application, tracks token spend as it happens, and attributes every dollar to a cost center. TrueFoundry, for instance, lets platform teams set hard caps so a single bad deploy can’t drain the AI budget. The detail matters because cost control only works if it’s enforced before the spend, not discovered after it. Security and credential sprawl Without a gateway, provider keys end up hardcoded in notebooks, committed to repos, and copied onto laptops. There’s no clean way to rotate t

TrueFoundry 2026-06-30 14:33 6 原文
AI 资讯 Dev.to

🚦 Meet Kueue: Smart Job Queueing for Kubernetes 🧠⚙️

Hey everyone 👋 If you run batch jobs, data pipelines, or any kind of AI and ML training on Kubernetes, you have probably hit this wall. Kubernetes is fantastic at deciding WHERE a pod should run, but it is surprisingly clueless about WHEN a job should start. 😅 You submit ten jobs, the cluster fills up, and the rest just sit there as Pending. No real queue, no priority, no fairness between teams. One noisy team can eat all your expensive nodes while everyone else waits. 🥲 That is exactly the gap Kueue fills, and today I want to walk you through it with a pile of hands on examples you can run on any cluster, even your homelab. 🏡 👉 Key takeaway up front: Kueue is a job level manager that holds your jobs in a real queue and only admits them when there is enough quota to actually run them. 🧪 Everything in this guide was tested against Kueue v0.18.1 using the v1beta2 API. I pinned every command and manifest to that version so you do not get surprised by API drift. 📋 What we will cover ✅ Why Kubernetes needs a queue ✅ The building blocks in plain language ✅ Installing Kueue ✅ Setting up quota with a ResourceFlavor, a ClusterQueue, and a LocalQueue ✅ Submitting a Job and watching it get queued and admitted ✅ Priority based admission ✅ Partial admission and elastic jobs ✅ Multiple resource flavors for x86 and arm ✅ Fair sharing between teams with cohorts ✅ Dedicated quota with a shared fallback ✅ Queueing a plain Pod ✅ Why this matters a lot for GPUs and your cloud bill 🤔 Why Kubernetes needs a queue Native Kubernetes scheduling is pod centric. The scheduler looks at one pod at a time and tries to place it. That works great for long running services. Batch workloads are different. They have a beginning and an end, they often need a fixed chunk of capacity, and they compete with other teams for the same nodes. Without a queueing layer you get: ✅ Jobs that fail or stay Pending when resources are tight ✅ No quota governance, so one team can starve the others ✅ No admission prio

Hamdi (KHELIL) LION 2026-06-30 14:32 11 原文
AI 资讯 Dev.to

The Illusion of the Clean Slate

Every engineer has fantasized about it: starting over. Throwing out the old system and building something clean. No legacy constraints. No accumulated compromises. Just pure, intentional design. It never works that way. You can delete all the code. You can architect from scratch. You can make the best technical decisions possible. But you can't delete the organizational memory. You can't unlearn what the last system taught you. You can't escape the patterns that already run through the business, the workflows people have shaped themselves around, the problems you've already paid the cost of understanding. The new system will look clean. But it will be haunted. What rewrites actually inherit A rewrite isn't a fresh start. It's archaeology pretending to be innovation. The constraints don't go away. The old system wasn't overcomplicated because engineers were bad. It was overcomplicated because of customer requirements, regulatory expectations, performance demands, and edge cases that took years to discover. A fresh rewrite finds all those edge cases again. Slower this time, because you don't have documentation—you have broken customers and escalations. The system gets layers of protection again, but now it looks like paranoia instead of learned caution. The organizational memory becomes invisible. Someone fought for that data model three years ago. There was a reason. A business rule that couldn't be violated. A data consistency requirement that cost a quarter to figure out. The new system doesn't have the battle scars that explain why things are the way they are. So they get rebuilt differently, until they hit the same requirement at 2am on a Saturday. The workflow is already baked in. Users have shaped their behavior around the old system. Sales has built their pitch around certain capabilities. Support has written documentation and runbooks. Customers have automation that depends on specific behaviors. The new system is technically cleaner, but it forces change on

Adam - The Developer 2026-06-30 14:28 6 原文
AI 资讯 Dev.to

A Prompt Is a Wish. A Tool Is a Law.

How I let non-engineers ship AI tools to production — and the boring infrastructure that made it safe. A product manager described a workflow in plain English — "every morning, pull yesterday's failed payments, group them by error code, and post a summary to our channel." Twenty minutes later it was running in production. She never opened an editor. She never saw a line of TypeScript. She talked to an agent, the agent wrote the code, and — once a human had reviewed the pull request — it shipped. That sentence should make you nervous. It made me nervous, and I'm the one who built the thing. The demo is "look, it wrote the code." The operation is "a marketer's tool now has a path to the payments database and nobody reviewed it." The interesting engineering isn't the part where an LLM writes code — that's the easy, demo-able part. It's the guardrails that decide whether the code it writes is allowed to exist. Here's the platform, and the five problems I had to solve to make it safe to hand to people who can't read the code that runs. The shape of the thing The platform is a place where anyone — engineers, PMs, designers, QA — can publish a reusable AI tool, and everyone else can use it. Write once, available to all. A few terms up front, because the whole design leans on them: MCP (Model Context Protocol) is a standard way for an AI client to discover and call your functions. The key detail: there's a step where the client asks the server "what tools do you have?" and the server answers with a list. Hold onto that — half the design hangs off that one list. Cloudflare Workers is code that runs on Cloudflare's servers at the network edge instead of your own. Durable Objects is per-session server-side storage that lives outside the model's context — the finite, token-costing window of everything the model can currently see. None of this is exotic; what matters is where each piece of state lives. Under the hood it's three small Workers speaking MCP: a gateway (auth, routin

Olexandr Uvarov 2026-06-30 14:25 7 原文
AI 资讯 Dev.to

AI Chunking Changes How We Should Build Content Pages

Traditional content pages are often designed for a linear reader. The introduction sets context, the middle develops the idea, and the conclusion ties everything together. AI retrieval does not always work that way. A system may identify smaller content units, pull the most relevant section, compare it with other sources, and use that fragment to support an answer. The full page still matters, but the retrievable blocks inside the page matter just as much. A useful Tumblr post explains the idea in simple terms: https://www.tumblr.com/digitalisedsoul/820825642809573376/ai-does-not-read-your-content-like-a-human?source=share For Dev Community readers, the pattern is familiar. Poorly structured inputs lead to weaker outputs. If content is dense, vague, or dependent on surrounding paragraphs, it becomes harder to extract and reuse. If content is modular, clear, and properly scoped, retrieval becomes easier. Marketing teams can learn a lot from this. A strong content page should behave like a set of well labelled components. Each section should answer a specific question. Headings should be descriptive, not decorative. Paragraphs should avoid vague references such as the above point or this approach when the section may be read independently. Definitions should appear close to the terms they explain. Examples should include enough context to stand alone. Proof should be written as text, not only displayed as graphics. Internal links should connect related concepts in a way that helps both readers and systems understand the topic map. A page about AI search visibility, for example, should not only include one broad explanation. It should break the topic into useful blocks: what AI visibility means, why AI systems retrieve passages, how source trust works, what makes content reusable, and how brands should measure answer presence. Each block becomes a possible answer unit. That structure does not weaken the reader experience. It improves it. Developers, marketers, and busi

harini work 2026-06-30 14:24 8 原文
AI 资讯 Reddit r/programming

Learning Neural Networking and making a one of my own

hey guys i'm a student of class 12 not expert but just interested to learn neural networking as today is my day second so i think to post my progress with you all as i can find the answer of question that i'm facing with. yesterday i have learn the basic of a neuron z=∑(w ⋅x )+b and today i'm learning about Batches, Layers, and Objects. despite =this it's hard to know what weight and bias really are if someone can explain me what they are please explain. that for today (>__<) submitted by /u/Vegetable_Cry_854 [link] [留言]

/u/Vegetable_Cry_854 2026-06-30 13:48 7 原文
AI 资讯 Product Hunt

Livinity

Open-source homeserver OS with a built-in AI agent Discussion | Link

Livinity IO 2026-06-30 12:36 4 原文