今日已更新 84 条资讯 | 累计 37333 条内容
关于我们

标签:#ratelimit

找到 5 篇相关文章

AI 资讯

The Subscription Squeeze: A Fortnight of Paying-User Gripes

Some fortnights the gripes scatter; this one they converged. Across the forums where paying customers of the big AI tools compare notes, the same complaint surfaced against three different companies in the same window: the monthly plan buys less than it did, and nobody dropped the price to match. One firm’s temporary generosity is about to expire, another’s year-old plan has quietly tightened, and the users caught in the middle are doing the same grim sum and reaching for the same coping strategies. Quotes sourced from: Reddit — specifically the subreddits r/ClaudeCode, r/perplexity_ai and r/Anthropic. Every quote below was opened at its permalink and copied verbatim; each is listed with its handle, subreddit and date in the Sources section. We quote experiences, not verdicts — a forum post is one person’s felt reality, and we have framed it as exactly that. The limit that lapses on the 19th The loudest note this fortnight came from Claude Code users watching a date on the calendar. Anthropic had lifted weekly limits by 50% as a promotion, extended it several times, and set it to expire around 19 August — after which allowances fall back to standard. For anyone who had adjusted their workflow to the higher ceiling, the lapse reads as a cut. A user posting as EnthusiasmMountain10 laid out the worry on 14 August: “I’m seeing more tokens burned on tasks that previously felt straightforward, more meandering, and generally less output per unit of usage. At the same time, the 50% usage reduction coming on Aug 19 makes this particularly concerning.” The complaint was double-barrelled: not only is the ceiling dropping, but the same work seems to cost more against it than it used to. That second half — a sense that quality had slipped — ran through the thread. A commenter posting as Captain_Birb put it bluntly: “quality is down.. 4.6 and 4.8 were sharper. Sonnet 5 is seriously a joke — it gets roasted by Opus every time and admits his faults.” Another, TheSassyPlant , descri

2026-08-23 原文 →
AI 资讯

OpenAI Is Testing a Button to Reset ChatGPT’s Limits — For $8

OpenAI is quietly testing a feature that lets ChatGPT users pay to undo their own usage limits. Hit the weekly cap on a $20 Plus plan and, for some users, a prompt now appears offering to restore the allowance to full for roughly $8. On the $200 Pro plan, the equivalent reset is reported to run up to about $80. The company never announced it; it was discovered by a subscriber who ran into it at the point of being locked out, and an OpenAI spokesperson later confirmed the company is exploring ways for capped users to buy more usage . The answer-first version: your flat monthly subscription now has a pay-to-continue button, and it shows up at the worst possible moment. The reset restores your usage to 100% and pushes the next weekly renewal about seven days out. It is cheaper than upgrading, which is the point — but it is also a new charge that did not exist a month ago, applied to a limit most users cannot see coming, offered at the instant they are least able to say no. What OpenAI is actually testing The mechanics, as reported, are straightforward. When a ChatGPT Plus subscriber exhausts their weekly message allowance, instead of only being told to wait, some accounts now see an option to pay to reset. Redeeming it restores usage to full and resets the weekly clock. The price sits at around $8 for Plus; on Pro it scales up to roughly $80, still framed as a stopgap against a full plan change. The feature was first surfaced by a Reddit user on the $20 plan who described a black prompt appearing at login once their allowance ran dry — not a setting they went looking for, but one that found them. That detail — a user, not a press release, breaking the news of a paid feature — is itself worth noting: the first public account of how OpenAI plans to charge for extra usage came from someone who had already been charged the inconvenience of being locked out. OpenAI has not disputed the reports. A spokesperson described the effort as exploring ways for people who exhaust the

2026-08-21 原文 →
AI 资讯

The Mark and the Meter: A Fortnight of AI Gripes

Some fortnights the complaints scatter across a dozen topics; this one they clustered around two. Across the forums where paying AI users gather to compare notes, two things happened in the same window that left the same taste: the meter moved, and the mark landed. In both cases, something about the product changed without a clear announcement, and the users found out by running into it. Quotes sourced from: the OpenAI Developer Community and Hacker News. Every quote below was opened at its permalink and copied verbatim; each is listed with its handle, platform and date in the Sources section. We quote experiences, not verdicts — a forum post is one person’s felt reality, and we have framed it as exactly that. The $200 plan that lasted two days The loudest note this fortnight was the Codex usage drain. On the OpenAI Developer Community, a Pro subscriber posting as jjjnoronha laid out the damage plainly on 4 August: “I’m on the $200 Pro plan, and my weekly limit was completely exhausted in 2 days of serious development work. This is simply not viable.” He was refactoring a Rust backend with two agents running — one coding, one checking code quality — and described leaving Ultra mode on for three or four hours, which alone consumed 30% of his weekly allowance. But the complaint was not that the work was expensive; it was that the plan was sold with language suggesting far more headroom than what the cap delivered. “When someone pays for the highest consumer tier,” he wrote, “the expectation is capacity for sustained, professional work — not being locked out for the rest of the week after two days.” Others described something stranger. A user posting as kalani reported going to bed with 70% remaining after a Saturday reset, then waking up to find 0%. “Not like I had something running overnight, and yet all usage blocked until this next Saturday,” they wrote on 4 August. “Hard to trust a $200 / month service that is completely unpredictable about whether you’ll even be

2026-08-20 原文 →
AI 资讯

Impact of deployment topology on rate-limiting and trust proxy

The trust proxy setting is an important concept in backend development, especially when implementing rate-limiting in our APIs. But deciding its accurate value depends heavily on our deployment topology. When we deploy our application in production, the client may not talk directly to our backend. There may be 1 or more proxies in between who forward the request to the next proxy or the backend server. Those proxies can be Load balancers, API gateways, reverse proxy like nginx or any custom service. So effectively, our request has to do some 'hops' over these proxies to reach backend. When we implement rate limiting in app to prevent the DOS attack, we generally intend this rate limit on the basis of client IP address. And this works fine when client request reaches our backend directly. But when we have multi-hop architecture, the simple setup won't work as expected. Because the most recent IP will be of the proxy and not the client. So all the traffic coming from different users will be considered from the single client(our own proxy) and thus there will be false positives as the rate limiting will trigger much often. In this scenario, we must tell our backend to ignore these extra hops(i.e. to trust our proxies). This is done by specifying trust proxy. If there is 1 proxy between client-server we set trust proxy to 1; if there are 2, or more, we set it accordingly. This will ensure our express app skips(trusts) these IPs, and accurately figures out actual client IP. The originating IP address of client is identified from 'X-Forwarded-For' header by the express app. But setting trust proxy is not that straightforward. The numerical value for trust proxy will not work in every case. If there are different paths from which our request reaches backend, there is a chance that the number of proxies may be different in each path. For example - internal vs external traffic: External(public) traffic: (Client -> Web Application Firewall -> Load balancer -> Reverse Proxy ->

2026-07-18 原文 →
AI 资讯

Rate Limiting — Throttling

Throttling: vì sao in-memory rate limit "biến mất" sau khi scale ngang, và chọn token bucket hay sliding window Throttling là cơ chế giới hạn số request một client (user, IP, API key, tenant) được xử lý trong một khoảng thời gian, để chống abuse, bảo vệ downstream, và phân bổ công bằng dung lượng service. Định nghĩa nghe đơn giản, nhưng lý do dev gặp nó trong việc thật lại rất cụ thể: sau khi scale service từ 1 pod lên 8 pod, cùng cấu hình "100 req/min mỗi user" đột nhiên trở thành 800 req/min thực tế — vì mỗi pod đếm riêng trong RAM, và load balancer rải request đều tám hướng. Rate limit vẫn "chạy", log không có lỗi, nhưng downstream vẫn bị flood. Đó là failure mode dẫn tới việc phải chuyển counter sang store phân tán, và kèm theo là câu hỏi chọn algorithm nào — token bucket, sliding window, hay leaky bucket — mỗi cái đánh đổi khác nhau. Cơ chế hoạt động Bốn thuật toán phổ biến, khác nhau ở cách đếm và cách xử lý burst. Fixed window counter. Chia thời gian thành khung cố định (mỗi phút bắt đầu tại giây 0). Mỗi request INCR một key rl:{user}:{minute} , nếu counter vượt limit thì reject. Đơn giản nhất, một INCR + EXPIRE trên Redis là xong. Nhược điểm cứng: tại biên khung có thể chịu gấp đôi limit trong một cửa sổ trượt — user gửi 100 req vào giây 59 của phút 12:00, rồi 100 req vào giây 01 của phút 12:01, tức 200 req trong 2 giây thật, trong khi limit là 100/phút. Sliding window log. Lưu timestamp của từng request trong sorted set, mỗi request ZADD + ZREMRANGEBYSCORE xoá các entry cũ hơn now - window , rồi ZCARD để đếm. Chính xác tuyệt đối nhưng tốn bộ nhớ tuyến tính theo số request. Sliding window counter. Cách Cloudflare mô tả trên engineering blog: giữ counter của khung hiện tại và khung trước, ước lượng lượng request trong cửa sổ trượt bằng nội suy có trọng số theo phần trăm khung trước còn nằm trong window. Chỉ tốn hai counter, sai số rất nhỏ so với log thuần, và không có failure mode biên như fixed window. Token bucket. Bucket có capacity B token, refill với tốc

2026-07-08 原文 →