标签:#r
找到 34027 篇相关文章
Odin, Wikipedia and Engagement Farming
Security guard, 72, behind design of Nike's new Shinjuku store
The UK's Latest "Debanking" Scandal Should Give Everyone Pause
The circuit that lets your brain think and see
Ants: Who looks after the injured in a colony?
Amsterdam invented the fire department
Giant trees have no trouble pumping water to top branches
Steam Controller Auto-Charge – pilot to magnetic charging puck using CV
Save Claude Code Tokens with Smart Routing
Dispersion loss counteracts embedding condensation in small language models
David Potter, the man who put Psion in the palm of your hand, logs off at 82
GitFut – Your GitHub stats turned into a World-Cup-style player card
Why Vancouver is always a stand-in for San Francisco in movies and TV shows (2021)
The Demoralization of the White-Collar Worker
GLM5.2 on AMD MI355X at 2626 tok/s/node at over 2x lower cost than Blackwell
EVE Online's Carbon engine is now open source: Fenris Creations explains why
Goodbye, Forever, Probably
Bought an expired domain. Then I inherited their AWS Root account
How We Vectorize 33.7M Ukrainian Court Decisions via Voyage AI
EDRSR — the Unified State Register of Court Decisions — is effectively all of Ukraine's judicial practice in open access. Today Qdrant holds **44M+ vectors : criminal (19M), civil (14.3M), commercial (5.1M), misdemeanors (5.6M). Vectorization of civil cases (CPC, justice_kind=1) — the largest cohort at 33.7M documents — runs on a dedicated EC2 instance (r6a.xlarge, 32 GB RAM, 2 TB gp3). Here's what's under the hood: models, pipeline, cost, rakes, and current status. Why Vectorize Courts When a lawyer searches "is there case law on recovering bank prepayment fees" — they don't want to open 40 decisions and read them through. They want the system to surface the top 5 most relevant ones, pull out key paragraphs, and show how courts reasoned. Full-text search (FTS) over keywords doesn't give that — it returns every document containing the word "fee", and there are thousands. For this semantic task you need vector representations of text. The model turns a paragraph from a decision into a point in a 1024-dimensional space; semantically similar paragraphs sit near each other. A kNN search in Qdrant returns the top K nearest, and an LLM composes the answer from exactly those relevant fragments. The only problem: the register is big. Very big. Scale Our prod database holds full texts of decisions starting from 2006. Breakdown by procedural type: Civil (CPC) — 33.7M documents. The largest category. Consumer, housing, labor, family. Criminal (CrPC) — 12M+ Administrative (CAS) — 14M+ Commercial (CC) — 6M+ Misdemeanors (CUaP) — 6M+ The Qdrant collection edrsr_decisions on a dedicated EC2 currently holds 44M+ vectors (122 segments, on_disk=true): | Proceeding type | justice_kind | Vectors | |—|—|—| | Criminal (CrPC) | 2 | 19,036,347 | | Civil (CPC) | 1 | 14,328,427 | | Misdemeanors (CUaP) | 5 | 5,579,432 | | Commercial (CC) | 3 | 5,098,662 | | Total | | 44,042,868 | Civil cases processed: 14.3M out of 33.7M — that's 42%. After CPC completes there will be roughly 63M+ vectors in