There are multiple models competing with 5.6sol on AA but none of them have the same feel (intuition,taste,judgement) - actually they are very far behind. I would say open source models are farther behind the the big labs than the benchmarks make you believe.
For simple edits it produces huge reasoning traces. Kind of disappointing. Super long repetitive reasoning always sits wrong with me. It feels like a way for the labs to brute-force higher benchmark scores but not actually like a smarter model. Disappointing, Mimo2.5 was such a nice model.
Steps 1–5 of his framework are increasingly formal ways of saying “figure out what’s going on before doing something,” followed by step 6: “then do it fast”
Seems like LLM can do everything Jev can do (just structured outputs?) but Jev is highly optimized and purpose built for it and thus way faster and cheaper. Is that a fair description?
I find it very intriguing. I reminds me of the insight in the early days that prompting the models to "think step by step" gave big performance boosts. Here we basically tell the models "You have an internal monologue/J-space and you can use it!". Quasi an opaque 'think step by step' prompt. The models were always able to think step by step but needed to be prompted to do so. Maybe the same thing is possible with the J-Space? The huge claims are yet to be replicated/proven, though.
Great thread, I was just thinking about compaction. My current line of thought is that compaction/pruning/ctx management in general should be something ongoing and maybe recursive. For example:
User:'How is auth implemented?'
->
[thinking]
[codebase exploration with [thinking] in between, 10 file reads, 3 of which were "wrong"]
[thinking]
->
agent_response
This little exchange contains a WHAT (how auth actually is implemented) and a HOW (where that info is and how to retrieve it). Maybe this question was part of a larger task. I think that whole exchange could be summarised before it enters context, kind of like what happens with subagents. The main thread would then consist mostly of [summaries]. Eventually the context will fill up anyway and we would summarise those summaries again. Alternatively one could maintain a [master_summary], kind of like an internal state. So new [summaries] get integrated directly and the [master_summary] gets updated.
reply