AI validated every real research hunch I fed it, and invented none of its own, across six months of actual forensic and steganography detection work — that's the case study, based on my own findings, not just a statement.
Every genuinely new direction in that work started as a hunch I couldn't fully explain yet. AI validated the hunch after the fact, showed it was real and measurable. It never generated the hunch itself. A 2024 Stanford study of 100+ NLP researchers backs this shape of result: LLM-generated research ideas scored higher on novelty in blind review, lower on feasibility — and when the researchers actually built the ideas out the following year, most of that novelty edge disappeared.
On pure brute-force parameter search: by round five of iterating a lead, the model itself starts suggesting we park it. Swapping to a more capable model doesn't change that. There's a documented way past it, though — not a smarter model, a verifier plus a search over thousands of candidates instead of one conversation at a time. FunSearch and AlphaEvolve do this at DeepMind's scale. More useful for a small team: OpenEvolve, a solo-engineer reimplementation with its own real, independently published results — a circle-packing result within 0.04% of DeepMind's own figure, and a documented GPU-kernel decode-speed improvement.
Not a clean win for AI, though. A January 2026 paper (Trehan & Chopra, arXiv:2601.03315) ran four fully autonomous ML-research attempts through a six-agent pipeline — three failed, for reasons that read like a direct description of the round-five wall: defaulting to training-data patterns, declaring success despite an obvious failure, weak judgement on what's even worth checking next.
So where does that leave the original claim? I don't think AI is quietly building its own successor in a loop that runs with nobody in the chain. Everything I've actually seen it do well — the reading, the validation, even the structured search that gets past brute force — still needs a person to define the problem first: what to search, what counts as a good result, when the problem itself needs rethinking. A loop that evolves entirely without humans would need AI to do that last part on its own. Six months in, I haven't seen it do that once.
That leaves me with a bigger question than the one I started with. Is a novel idea, real innovation, actually tied to imagination — the ability to dream something up, or that inkling you get when you suspect there's a direction worth taking before it even makes sense to explain why? That feels like a genuinely human thing. Will AI ever endeavour into that space on its own terms, or will it keep covering the same ground through brute force — generating enough new-looking concepts that it eventually stumbles onto what a person would have reached by instinct? Maybe, in time. I don't know yet.
If you've found a real way to get AI past brute force in a small-team setting, not a DeepMind-scale one, I'd like to hear how.




