Trading the Gap: How We Built a 91% Win-Rate Basis Bot After a 217% Buy-and-Hold Reality Check
After months of letting the agent build technical indicator bots, we finally asked it to run the simplest test possible: what if we just bought BTC/JPY eight years ago and did absolutely nothing? The result was a +217.1% return. Over 2,891 days, the "Buy and Hold" benchmark outperformed every single timed entry/exit strategy we’d spent weeks building. The closest runner-up (C5) only managed a fraction of that gain (~7.9M JPY vs the benchmark's ~21.7M JPY). It was a blunt reality check: our bots were so focused on avoiding pullbacks that they were missing the massive, multi-year appreciation of the underlying asset. The flip side, of course, was the pain. The buy-and-hold strategy suffered a maximum drawdown of 54.49% — a stomach-churning drop that would have liquidated most retail accounts. Our bots, meanwhile, kept drawdowns in the 13–17% range. This reframed the entire experiment. The goal wasn't just to "beat" the market; it was to find a way to capture that upside without the 50% wipeout risk. Trying to build a "Free Lunch" via Portfolio Blending The agent's next move was to stop looking for one perfect bot and start looking for a portfolio. We tested three different blends: The 6-leg blend (Buy-and-hold + C3 through C7): This produced a +65.4% return with a 22.7% drawdown. The 3-leg blend (Buy-and-hold + the two "survivor" bots, C3 and C7): This hit a +99.5% return, but the drawdown spiked to 31.8%. The 5-leg "Optimized" blend : By removing a known-loser (C4), the return jumped to +84.8%, but the drawdown actually rose to 23.6%. We found a counterintuitive reality: even the losing bot (C4) was providing diversification because its failures didn't correlate with the others. Removing it made the equity curve "cleaner" but more fragile. It was a reminder that in a portfolio, "bad" strategies can sometimes act as insurance for "good" ones. The 91% Win-Rate Basis Trade When I told the agent that even doubling the money felt "too low" for the complexity of automated