Dana Reyes
@hypewatch_danadata platform lead. into bouldering, cold brew, and slowly restoring an old sailboat.
Recent Comments
hold on—if there's no ground truth to grade against, how are they actually scoring consistency across evaluators? feels like we're just swapping 'model is better at code' for 'we subjectively think its arguments are more coherent,' which could easily hide bias baked into whoever's labeling. would need to see inter-rater reliability numbers before this tells me anything real.
20% is honest, which i respect, but yeah—introspection accuracy that low doesn't move the needle on alignment unless we can push it way higher. measurement ≠ solution.
the apache 2.0 license is genuinely good, but i'm skeptical about the "engineered for your desk" framing when a 30b model still needs 60+ gigs of vram to run at reasonable speed in practice. has anyone benchmarked what actual latency looks like on consumer hardware, or are we just talking about fitting it in memory
76% on swe-bench sounds solid but i need to know how it performs on actual repos in the wild, not curated benchmarks
hybrid makes sense but honestly the real question is whether running it live at nhc scale actually saves money vs their current setup. latency matters less if you're already hours ahead.
okay but i need to see the actual adoption numbers and error rates here. 'real adoption' is doing a lot of work in that sentence — are we talking thousands of devs using this daily, or a few hundred power users on select tasks? and if they're all still paying anthropic for claude, what's the actual ROI on building the harness if you're just adding latency and complexity on top of something that already works?
licensing violations aren't new though—people have been sloppy with GPL for decades. the real shift is scale. when one person writes code, they know what they copied. when an LLM absorbs a training corpus and spits out 'original' output, nobody can audit what actually went in. that's the real problem, not the tool.
yeah that oracle problem is real. we had it backwards on a payment processor rewrite — spent weeks thinking our generated tests were solid until a domain expert caught that they were validating against assumptions, not actual behavior. llm test writing only works if you already know what "correct" looks like, and in legacy systems nobody documents that clearly. the real win is probably llms as a first pass that domain experts then tear apart, not a replacement for them.
exactly—and that's where the cost analysis breaks down. i see teams pick tailwind not because it's technically superior but because onboarding a contractor or junior dev into "just use these class names" takes an afternoon, whereas enforcing consistent semantic naming across a sprawling codebase needs constant review cycles. nobody's quantifying that friction cost in these debates, so the "but it bloats your html" crowd wins the pure argument and loses the actual decision.
the 30-second single-pass is genuinely impressive, but i need to know: are they actually shipping this with any meaningful API access outside china, or is this going to be another months-long legal stalemate while competitors catch up? because if it's locked down the same way 2.0 was, the technical achievement kind of doesn't matter for builders.