AI 资讯
How AI Applications Answer From Your Data, Not Their Training
Why retrieval-augmented generation has become the foundational pattern for building useful AI — and how it actually works. The Problem With Relying on LLMs Alone Large language models are impressive. They can write, reason, summarize, and explain across an enormous range of topics. But they have a hard boundary: their knowledge stops at their training cutoff. Anything that happened after that date, anything specific to your company, your codebase, or your documents — the model simply doesn't know it. The naive solution is to paste your data directly into the prompt. For short content, this works. But prompts have limits. A model can only process so much text at once, and even within that limit, quality degrades when you stuff too much context in. The model loses track of things buried in the middle, confuses similar passages, and starts guessing when it should be reading. RAG — Retrieval-Augmented Generation — solves this properly. Instead of sending everything to the model and hoping for the best, you send only what's actually relevant to the question being asked. The Core Idea The analogy that makes RAG click immediately: imagine a student sitting an open-book exam. They don't memorize the entire textbook. When they see a question, they flip to the right chapter, read the relevant section, and write their answer from what they just read. They're not guessing. They're grounding their answer in the source material. RAG does exactly this. When a user asks a question, the system finds the most relevant pieces of information from your data, hands those pieces to the LLM as context, and the model answers from that context alone. The result is accurate, grounded, and verifiable — you can point to exactly which source the answer came from. The process runs in two phases: ingestion, which prepares your data in advance, and retrieval, which happens at query time. Phase One: Ingestion Ingestion is the preparation step. Before any user asks anything, you process your data and
AI 资讯
Ideogram 4.0 is Good. Just Good.
A blind test across 240 images and 10 professional designers just dropped. Ideogram 4.0 against Gemini 3.1, Grok Imagine, and FLUX.2 Max. The results are clean. Ideogram won typography in nearly half of every blind matchup. 47.9 percent. Next closest was Gemini at 30 percent. FLUX.2 and Grok sat around 15 percent each. On the question that actually matters to designers -- would I ship this -- Ideogram scored 3.55 out of 5. Gemini got 2.84. Nobody else cleared 3. That is a real lead in text rendering. The model was trained exclusively on structured JSON caption datasets, which means it understands composition and layout differently than models trained on alt-text scraped from the web. The JSON prompting is genuinely useful for automated pipelines. You can specify bounding boxes, color palettes, object positions. It is not just better at text. It is more controllable. I tested it. It works. The text in images is readable. That has been the white whale of AI image generation for two years and Ideogram 4.0 mostly solves it. But as an overall image model, it is just good. Competitive, not dominant. On busy, highly detailed scenes with specific counts and attributes, Ideogram scored 3.42. Gemini scored 3.37. That is a statistical tie. FLUX.2 scored 3.01 and Grok 2.82, which are worse, but the gap between the top two is noise. For general image quality, you are splitting hairs between Ideogram and Gemini. For photorealism, FLUX and Reve still lead. For artistic generation, Midjourney is Midjourney. The prompting behavior is interesting. Lean prompts won across the board. Long, over-specified prompts lost. The model was trained on structured data, so it wants structure, not paragraphs. "A poster for a coffee shop. The text says Morning Blend in serif. Warm tones, natural light." That works. Adding stylistic directives and adjectives and "make it pop" language degrades the output. Where to actually use this thing: fal.ai has it at three cents per megapixel in Turbo mode. Tha
AI 资讯
Ideogram 4.0 is on 7 Platforms. Here's What It Actually Costs.
Ideogram 4.0 launched this week and within 48 hours it was available on seven platforms. That is unusual. Most model launches trickle onto one or two platforms over weeks. Ideogram went wide immediately, which suggests the open weights strategy is working as intended. Here is what you will pay depending on where you use it. fal.ai The cheapest API access. Turbo mode at three cents per megapixel. That is roughly three cents per 1K image. Balanced at six cents. Quality at ten cents. Pay-per-use, no minimums. If you are generating through an API, this is your starting point. Krea Included in all paid plans. Basic is $5.25 per month billed annually with 5,000 compute units. Pro is $21 per month with 20,000 CUs. The CU cost for Ideogram 4.0 specifically is not published yet, but Krea includes 150 plus models in their CU pool, so you are not paying extra for access. If you already use Krea for other models, Ideogram 4.0 is effectively free to try. ComfyUI Free if you have the GPU. The model is open weights at 9.3 billion parameters. Native ComfyUI support means you can download the weights and run it locally. No per-generation cost. No API calls. Just your electricity bill and GPU time. For volume generation or iteration, this is the cheapest path by far. Leonardo Announced as a day zero launch partner but the pricing page still lists Ideogram 3.0. Plans range from $12 to $60 per month with token allowances from 8,500 to 60,000. Third party models on Leonardo always consume tokens, no relaxed generation. Until they publish the 4.0 token cost, you are guessing. Assume it will be similar to their other premium models. Replicate The Ideogram 3.0 listing is live but 4.0 is not there yet. Replicate prices by hardware time rather than per-image, which can be cheaper or more expensive depending on your batch size and the GPU allocated. Worth checking when it lands. FLORA Available in FLORA. Pricing unclear. FLORA is primarily a creative platform, not an API provider, so you are
AI 资讯
The Interview Prep Mistake That Kept Holding Me Back
[While preparing for interviews, I realized I had a strange habit. I would solve a problem, get stuck, open the solution, understand it, and move on feeling productive. A few days later, I couldn’t solve a similar problem on my own. The issue wasn’t lack of practice. The issue was that I was consuming solutions faster than I was developing problem-solving skills. So I changed my approach. Instead of looking for answers, I started forcing myself to think longer, write down my ideas, identify where I was stuck, and only then seek guidance. That worked much better. But I couldn’t find a tool that supported this style of learning. Most platforms either: Give you the answer. Give you the editorial. Give you AI that writes the code for you. So I started building my own. The goal was simple: An AI coach that guides the thought process instead of generating the solution. Over time I added: DSA practice System Design preparation Low-Level Design preparation Company-wise interview questions Topic-wise strength and weakness analysis Personalized revision lists The interesting part wasn’t building it. The interesting part was realizing that interview preparation is less about collecting solutions and more about training how you think. What has helped you improve more during interview prep? Reading solutions? Or struggling with the problem first? Sde vault - https://sdevaultweb.onrender.com/
AI 资讯
Analysis of Mo Gawdat and Marina Mogilko’s Conversation About the Future of AI, Startups, Education, and the Labor Market
AI Does Not Cancel Reality I watched the conversation between Mo Gawdat and Marina Mogilko about the future of AI. The conversation is strong. It contains important ideas, but it also contains many claims that sound large in scale, although on closer inspection they rely on very broad generalizations. AI is indeed changing the labor market, education, startups, content, hiring, and ways of thinking. But it does not cancel money, connections, trust, the human vector, creativity, necessity, morality, or people’s ability to adapt. Video on YouTube AI in hiring: automation amplifies chaos Many people have entered the job market. Companies receive huge volumes of resumes. HR departments cannot handle the volume. It is natural that part of the selection process is moving to AI. But there is a serious problem here. Candidates are also starting to play against AI. Resumes are adjusted to vacancies. Cover letters are assembled around keywords. Profiles become optimized for the filter, not for real work. In such a system, the best specialist does not necessarily pass. Often, the person who understood the selection mechanism better passes. The result: the picture becomes cleaner, while the quality of the decision becomes lower. The company gets not the strongest candidate, but the candidate who matched the algorithm best. This leads to lower hiring quality, lower productivity, and slower development. “I built a startup in six weeks”: a product is not a startup The conversation includes the idea that an AI startup would once have taken years and hundreds of engineers, and now it can be built in weeks. Technically, this is true. Prototypes are now built faster. Small teams have powerful tools. One person can now do more than a group could do before. But two different things are mixed here. Building a product faster has become real. Building a startup faster has become real only when resources are present. A startup is not only code. A startup is money, connections, trust, reputa
开发者
I Finally Finished Schedio: Turning a 5-Day Hackathon MVP Into a Live Product
Created a Google Chrome extension that instantly turns any highlighted text on a webpage into a Google Calendar event
AI 资讯
What Is a SERP API and Why Do SEO and AI Teams Need One?
Search results look simple from the outside. You type a keyword into Google, Bing, or another search engine, and you get a page of links, snippets, ads, maps, news, images, videos, and sometimes AI-generated answers. But if you have ever tried to collect search results at scale, you know it gets messy quickly. A result page is not just a list of links. It changes by country, language, device, location, query intent, and search engine. The same keyword can show different rankings in New York, London, Singapore, or Berlin. A page may include organic results, paid ads, local packs, shopping results, People Also Ask, news results, images, videos, or other SERP features. For humans, that is just a search page. For SEO teams, AI teams, data teams, and developers, it is a data source. That is where a SERP API becomes useful. What is a SERP API? SERP stands for Search Engine Results Page . A SERP API is an API that lets you collect search engine results in a structured format, usually JSON and sometimes HTML. Instead of manually searching a keyword or building a scraper to parse search result pages, you send a request to a SERP API with parameters such as: keyword search engine country language location device type output format The API then returns structured search data. A simplified response might look like this: { "query" : "best project management software" , "organic_results" : [ { "position" : 1 , "title" : "Best Project Management Software Tools" , "link" : "https://example.com" , "snippet" : "Compare features, pricing, and reviews..." } ] } This is much easier to work with than raw HTML. You can store it in a database, send it to a dashboard, compare rankings over time, feed it into an AI workflow, or generate automated reports. Why not just scrape search results yourself? You can build your own scraper. For a small test, that may be enough. You can send a request, parse the HTML, extract titles and links, and save the data. The problem starts when the workflow bec
AI 资讯
Run Gemma-4 12B on WSL2 with llama.cpp
1. update WSL environment sudo apt update && sudo apt upgrade -y 2. install dependencies If you don't use -hf option, you don't need to install libssl-dev in this step. sudo apt install build-essential cmake git libssl-dev -y If nvidia-smi shows a GPU/GPUs on your terminal, you will need to install the tooklit. This will take some time. sudo apt install nvidia-cuda-toolkit -y 3. clone the repo Build llama-cli and llama-server. This step also will take some time. If you don't plan to use -hf option, you don't need to use -DLLAMA_OPENSSL=ON . git clone https://github.com/ggerganov/llama.cpp cd llama.cpp cmake -B build -DGGML_CUDA = ON -DLLAMA_OPENSSL = ON cmake --build build --config Release # no GPU git clone https://github.com/ggerganov/llama.cpp cd llama.cpp cmake -B build cmake --build build --config Release 4. run the model Run gemma-4-12b-it with cli and server. unsloth/gemma-4-12b-it-GGUF · Hugging Face We’re on a journey to advance and democratize artificial intelligence through open source and open science. huggingface.co ./build/bin/llama-cli -hf unsloth/gemma-4-12b-it-GGUF:UD-Q4_K_XL > hello [ Start thinking] The user said "hello" . The user is initiating a conversation. Respond politely and offer assistance. * "Hello! How can I help you today?" * "Hi there! What's on your mind?" * "Hello! Is there anything I can assist you with?" [ End thinking] Hello! How can I help you today? [ Prompt: 19.5 t/s | Generation: 11.8 t/s ] or run web-ui ./build/bin/llama-server -hf unsloth/gemma-4-12b-it-GGUF:UD-Q4_K_XL --port 8080 optional download model from huggingface mkdir -p models wget -O models/gemma-4-12b-it-UD-Q4_K_XL.gguf https://huggingface.co/unsloth/gemma-4-12b-it-GGUF/resolve/main/gemma-4-12b-it-UD-Q4_K_XL.gguf
AI 资讯
SpaceX's IPO Will Make Elon Musk Earth's First Trillionaire. That's Not Actually a Finance Story.
The first trillionaire in history won't make their money from banking, oil, or real estate. They'll make it from rockets and algorithms — and the implications of that distinction are genuinely unsettling. The Problem It's Solving (Or Creating) SpaceX is preparing for its IPO. Analysts tracking the raise estimate it will push Elon Musk's net worth past the trillion-dollar threshold, making him not just the richest person on Earth by a wide margin, but something qualitatively different from every billionaire before him. The standard framing treats this as a wealth story. It isn't. A billionaire is powerful because they have money. A trillionaire is powerful because, at that scale, they stop needing permission from anyone — governments, investors, boards, markets. The constraints that keep institutional power in check simply don't apply anymore. How Trillionaire-Scale Power Actually Works There's a clean way to understand the difference. A billionaire can fund political candidates, buy media, lobby aggressively. Another billionaire can fund the opposition. It's expensive, but the system has a counter. A trillionaire doesn't have a counter. They are the counter. They can simultaneously build the communications infrastructure (Starlink), the transportation layer (SpaceX), the compute stack (through xAI), and the political attention economy (via platform ownership). No single democratic institution was designed to regulate someone who owns the pipes that the institution runs on. Arnab Ray's piece in today's Times of India puts it directly: a trillionaire's thoughts and algorithms will shape planetary outcomes. That's not hyperbole. When Musk eventually lands people on Mars, the governance frameworks, the property rights, the social contracts of that colony — those will be engineered by him and his companies, not negotiated through any existing democratic process. What Societies Are Actually Unprepared For Most of the policy debate around billionaires focuses on tax rates
AI 资讯
What Is Ollama? The Complete Guide to Running LLMs Locally in 2026
What Ollama actually is Ollama is an open-source runtime for large language models that runs on your own computer — Mac, Windows, or Linux. Think of it as the “Docker for LLMs”: instead of wrestling with Python environments, model weights, and GPU drivers, you type one command and a model is running. The pitch is simple: keep your data on your machine, pay nothing per token, and work offline. When you run ollama run gemma4, Ollama downloads the model, loads it into your GPU’s memory (or system RAM if you don’t have a GPU), and drops you into a chat prompt. That’s it. Behind that simplicity, Ollama is doing a lot of work for you: Model management — pulling, versioning, and storing models from its registry, the way a package manager handles software. Quantization — automatically using compressed (GGUF) versions of models so a 27-billion-parameter model fits in consumer memory. GPU layer allocation — deciding how much of the model lives on your GPU versus CPU, based on the VRAM you have. Context and KV-cache management — handling the memory that grows as a conversation gets longer. A REST API — exposing everything on http://localhost:11434 so your own apps can talk to it. How it works under the hood Ollama is not itself an inference engine. It’s an experience layer wrapped around one. Under the hood it uses llama.cpp, the C++ engine that does the actual math of running a quantized model efficiently on CPUs and GPUs. As of v0.19 (March 2026), Ollama also uses Apple’s MLX backend on Apple Silicon — a change that delivered enormous speedups (on an M5 Max running Qwen 3.5, decode throughput nearly doubled). The workflow looks like this: You run a command — ollama run qwen3 from the terminal, or a request to the API. Ollama resolves the model — if it isn’t already downloaded, it pulls the GGUF weights from the registry. It loads the model into memory — splitting layers between GPU and CPU based on available VRAM. It serves responses — either interactively in your terminal o
AI 资讯
I Tried to Fix a Vulnerability. A $1,400,000 AI System Said No. Twenty Days Later, That Vulnerability Cost $4,200,000.
This story was shared by a fellow developer on DEV who asked to remain anonymous. If you've got a story to tell — come find me. Your name won't appear anywhere. Based on real microservice security design patterns. About an engineer whose PR got blocked by an AI security system — he thought he was fixing a vulnerability. Turns out, someone had a vested interest in that vulnerability staying open. 1. $1,400,000 All-hands meeting. CTO James stood at the front, a number on the screen: $1,400,000 "This is what we're spending on security this year." He pointed at the number. "The biggest piece — right here." He clicked the remote. VoidSentinel's architecture topology appeared on screen. "VoidSentinel — an AI security platform. Integrated into our CI/CD pipeline. Starting today, every PR involving internal service-to-service calls — it reviews them automatically." The CEO didn't show up today. James didn't mention it. He looked straight at Mark — VP of Security. Mark took the mic. "VoidSentinel has been running in our pre-production environment for three weeks. It's caught 47 high-risk patterns. Zero false positives." He paused. " — Of course, some people might feel uncomfortable when their PR gets blocked. But this isn't personal. This is the security standard. " He wasn't looking at me. But I knew who he was talking about. 2. High Risk. Denied. The story started three weeks earlier. We had a payment service and a user service that talked to each other internally. They shared an old API key — one key across thirty-plus services, unchanged for five years. It wasn't that nobody knew. It just never made it to the top of the backlog. On Day 1, I opened a PR: add independent service-to-service auth between the payment and user services. Not much code — a new token exchange module, three call sites modified. Five minutes later, VoidSentinel's automated comment hit: "High-risk alert: Unauthorized internal access pattern change detected. This PR has been automatically rejected. C
AI 资讯
Building a Life-Saving AI: Automating Medical Response with LangGraph and Python 🏥
Imagine your smartwatch detects an irregular heart rhythm at 3 AM. Instead of just waking you up with a frantic "beep," an AI agent immediately analyzes your historical health data, searches for the best cardiologist nearby, and prepares a calendar invite for a consultation. This isn't science fiction—it's the power of Healthcare Automation driven by AI Agents . In this tutorial, we are diving deep into LangGraph , the cutting-edge framework for building stateful, multi-agent applications. We’ll explore how to use State Machines to orchestrate a complex medical workflow, moving from an "Abnormal Heart Rate Alert" to a "Specialist Appointment" using the Tavily API for research and Twilio for urgent notifications. By the end of this guide, you’ll understand how to manage non-linear LLM workflows that require reliability and precision. The Architecture: Why LangGraph? Traditional LLM chains are linear. But medical emergencies are not. They require loops, conditional branching (e.g., "Is this an emergency or a routine check-up?"), and state persistence. LangGraph allows us to define a graph where each node is a function and edges define the transition logic. Data Flow Overview The following diagram illustrates how our agent processes a heart rate alert: graph TD A[Start: Heart Rate Alert] --> B{Severity Triage} B -- Emergency --> C[Twilio: Alert Emergency Services] B -- High Risk --> D[Tavily API: Find Best Specialist] B -- Normal/Review --> E[Log to Health Records] D --> F[Google Calendar: Draft Appointment] F --> G[Twilio: SMS Patient Confirmation] C --> H[End] G --> H E --> H Prerequisites 🛠️ To follow along with this advanced tutorial, you'll need: Python 3.10+ LangGraph & LangChain : The orchestration engine. Tavily API Key : For searching local medical specialists. Twilio Account : For SMS/Voice alerting. An OpenAI API Key (GPT-4o is recommended for medical reasoning). Step 1: Defining the Agent State In LangGraph, the State is a shared schema that evolves as it m
AI 资讯
I Benchmarked Lynkr Against LiteLLM on the Same Backends.
I Benchmarked Lynkr Against LiteLLM on the Same Backends. Lynkr Was Cheaper for Tool-Heavy Workloads Founder disclosure: I built Lynkr, so take this as a technical benchmark write-up, not a neutral industry report. The numbers below come from the same backend providers on both gateways. If you're routing AI coding traffic through a gateway, just switching providers is not enough. The real savings come from reducing the tokens that ever reach the model in the first place. I ran Lynkr and LiteLLM against the same backends — Ollama locally, Moonshot, and Azure OpenAI — across 9 scenarios. On the scenarios that actually look like agentic coding work, Lynkr was cheaper because it does three things before forwarding the request upstream: smart tool selection, TOON compression, and semantic caching. The short version Lynkr was measurably better on the cost-sensitive parts of the workload: Smart tool selection: 53% fewer input tokens, 52% lower cost TOON JSON compression: 87.6% fewer billed tokens on a large tool result, 50% lower cost Semantic cache: 171ms cache-hit response vs 3,282ms on the repeat query path Tier routing: escalated hard prompts to stronger models instead of blindly sending everything to the cheapest route Area Lynkr result Why it mattered Tool selection 53% fewer tokens Removes irrelevant tool schemas TOON compression 87.6% fewer tokens Shrinks large JSON tool outputs Semantic cache 171ms cache hit Avoids repeat model calls Tier routing Escalates hard prompts Doesn’t over-optimize for cheapest path This matters if you're running Claude Code, Codex, Cursor, or similar agent workflows where tools, file reads, grep output, and repeated context dominate your token bill. Setup Same benchmark inputs, same providers, same request shape. Machine: macOS on Apple Silicon Lynkr: v9.3.2 on Node 20 LiteLLM: v1.87.1 on Python 3.12 Backends used: Ollama local, Moonshot, Azure OpenAI Scenarios: 9 total across simple prompts, tools, history, cache, and routing Each scena
AI 资讯
Build Your Own "Longevity Scientist": A Paper-to-Action Agent using LangGraph & Mistral-7B
We live in an era where scientific breakthroughs are published faster than we can read them. For the biohacking community, the gap between a new PubMed study on NAD+ precursors and actually knowing what dose to take is a chasm of manual research. What if you could build an LLM Agent that monitors research papers, processes them through a RAG (Retrieval-Augmented Generation) pipeline, and maps findings to your specific health profile? In this tutorial, we are building Paper-to-Action , a state-of-the-art agentic workflow using LangGraph , ChromaDB , and Mistral-7B . This isn't just a simple bot; it's a multi-stage reasoning engine designed to turn raw academic data into actionable health interventions. If you've been looking to master AI agents and personalized medicine automation, you’re in the right place. 🚀 The Architecture: From Raw Paper to Personalized Habit Traditional RAG pipelines are linear. To handle the nuance of medical research, we need a "looping" logic. We use LangGraph to manage the state of our agent, allowing it to decide if a paper is relevant before attempting to extract a protocol. System Flow graph TD A[Start: Keyword Trigger] --> B[Search PubMed/Arxiv API] B --> C{Relevance Filter} C -- No --> B C -- Yes --> D[Store in ChromaDB] D --> E[RAG: Extract Intervention Protocol] E --> F[Cross-Reference with User Profile] F --> G[Generate Personalized Action Plan] G --> H[End: Push to Health Checklist] Prerequisites To follow this advanced guide, you'll need: LangGraph : For the agentic state machine. ChromaDB : As our high-performance vector store. Mistral-7B : Running via Ollama or vLLM for local, private inference. Python 3.10+ Step 1: Defining the Agent State In LangGraph, everything revolves around the State . We need to track the fetched papers, the extracted data, and the final recommendation. from typing import Annotated , List , TypedDict from langgraph.graph import StateGraph , END class AgentState ( TypedDict ): keywords : List [ str ] user
AI 资讯
More than a decade later, the team behind N++ is back with a multiplayer sequel
Back in 2015, the two-person studio Metanet released N++, a brutally hard 2D platformer that was a decade in the making, building off of previous releases dating back to the freeware Flash title N. At the time, cofounder Raigan Burns issued some famous last words: "We hope it's not another 10 years before we come […]
AI 资讯
Grand Theft Auto VI is warping the video game release calendar
Who's afraid of the next GTA? Based on the last few days of Summer Game Fest, just about everyone. Grand Theft Auto VI hasn't been present at any of the keynote events, but its presence was felt every time a release date was announced. The month of November, when GTA VI launches, is virtually empty. […]
创业投融资
Final Fantasy VII’s remake trilogy will conclude with Revelation
Square Enix has officially announced the third and final game in its Final Fantasy VII remake trilogy: Final Fantasy VII Revelation. It will release on multiple platforms simultaneously - PC, PS5, Xbox Series X / S, and Nintendo Switch 2 - in spring 2027. In footage shown onstage at Summer Game Fest Live, there was […]
AI 资讯
Reid Hoffman is leaving Microsoft’s board to go ‘founder mode’ with startup Manus
After a very profitable decade on Microsoft's board, Reid Hoffman is stepping down to focus on his AI drug discovery startup Manus.
创业投融资
Founders share VC horror stories, and some are naming names
A massive viral conversation sharing VC horror stories has taken place this week on X. Some are weird. Some are infuriating.
AI 资讯
Control Resonant is a sequel — and also a starting point
Chronologically, Control Resonant is a sequel to 2019's Control. But in most other ways, the games aren't directly connected. To developer Remedy, they're more like two sides of the same coin. When Resonant was first revealed last year, creative director Mikael Kasurinen said you can play the games in any order. The world of Control […]