Yes it does. Just need to go beyond prompt. Add reinforcement loops, make it test the solution in a separate environ it can’t screw up in, and have it not just INVENT things (“give me X”), but research the topic and base its solution on the rules created by the research.
This is what basically the Claude harness (not the local but the remote harness you can’t see) adds to the LLM what makes it so powerful and useful. Replicate those processes and even a small 4B mode will be incredibly capable.
No amount of prompt “engineering” will help when they outright make shit up.
Yes it does. Just need to go beyond prompt. Add reinforcement loops, make it test the solution in a separate environ it can’t screw up in, and have it not just INVENT things (“give me X”), but research the topic and base its solution on the rules created by the research.
This is what basically the Claude harness (not the local but the remote harness you can’t see) adds to the LLM what makes it so powerful and useful. Replicate those processes and even a small 4B mode will be incredibly capable.