Yes, older models will make those mistakes. The solution is to provide appropriate agent/skill definitions so it doesn’t just spit something out, but rather comes up with the solution, then smoke tests it in a separate scratch environment.
Claude uses this approach for reinforced learning, and it works well. Takes a few more rounds to resolve, but the solution is generally flawless (for that specific purpose).
I use Claude + local models.
Yes, older models will make those mistakes. The solution is to provide appropriate agent/skill definitions so it doesn’t just spit something out, but rather comes up with the solution, then smoke tests it in a separate scratch environment.
Claude uses this approach for reinforced learning, and it works well. Takes a few more rounds to resolve, but the solution is generally flawless (for that specific purpose).