This is basically the same as asking someone who knows a bit about everything to recite obscure information from memory. I doubt most humans would do better.
What you’re missing here is a feedback loop, validation, skills, etc. For example, if I ask one of my well-configured agents the same question, they go read the help/man page/other docs, start a vim session, quickly test/iterate until the result is correct, then give me a one-page document explaining what I need to know, with cited evidence — far faster than I’d do it, and I can keep working for the few seconds it takes. This rigor is written into my global instructions, not something you get out of the box on most models (Anthropic’s models tend to be good at this without handholding, which is part of why they’re so popular, aside from the fact that they just don’t make as many mistakes).
The same mindset scales to larger software problems, too. As long as you have a well-defined specification and good agent instructions (and/or something like Spec Kit), you can have agents break it down, implement, and then other agents compare the result to the spec, and just keep looping until it is done. The hard part is writing good specs and requirements, but that’s not a new problem.
This is basically the same as asking someone who knows a bit about everything to recite obscure information from memory. I doubt most humans would do better.
Sure, but a good web resource or even a good reference book would provide better help faster. Unfortunately it’s getting harder to harder to find those good web resources as search results get overtaken by AI slop.
iterate until the result is correct, then give me a one-page document explaining what I need to know, with cited evidence — far faster than I’d do it
I have gotten it to do this in certain situations, like the other day I wanted to simplify a math formula, so I gave it a loop with a Python script that checked its solution against the original reference version. This worked pretty well.
But I had to write code specifically for that situation. Even with a good skeleton to start with, it’s a non-negligible amount of work to get that set up. I feel like the scenario in which this is useful is kinda narrow: when I have a very good idea of exactly what I want, but some step along the way is a hassle. General software engineering, like making a whole app, is far too open ended, and most of the sub-problems I encounter in software engineering seem either too open ended or too small to benefit from this approach.
That also doesn’t account for the speed of models. My experience is a ~20GB locally hosted model takes like 1-5 minutes to produce a good length response, and the few times I have used online models they are often slower. A few minutes per iteration, accounting for debugging when it goes off track, is not exactly fast or hassle free.
I just feel like the trade off where using AI vs doing it all myself is pretty limited in when AI offers an advantage.
This is basically the same as asking someone who knows a bit about everything to recite obscure information from memory. I doubt most humans would do better.
What you’re missing here is a feedback loop, validation, skills, etc. For example, if I ask one of my well-configured agents the same question, they go read the help/man page/other docs, start a vim session, quickly test/iterate until the result is correct, then give me a one-page document explaining what I need to know, with cited evidence — far faster than I’d do it, and I can keep working for the few seconds it takes. This rigor is written into my global instructions, not something you get out of the box on most models (Anthropic’s models tend to be good at this without handholding, which is part of why they’re so popular, aside from the fact that they just don’t make as many mistakes).
The same mindset scales to larger software problems, too. As long as you have a well-defined specification and good agent instructions (and/or something like Spec Kit), you can have agents break it down, implement, and then other agents compare the result to the spec, and just keep looping until it is done. The hard part is writing good specs and requirements, but that’s not a new problem.
Sure, but a good web resource or even a good reference book would provide better help faster. Unfortunately it’s getting harder to harder to find those good web resources as search results get overtaken by AI slop.
I have gotten it to do this in certain situations, like the other day I wanted to simplify a math formula, so I gave it a loop with a Python script that checked its solution against the original reference version. This worked pretty well.
But I had to write code specifically for that situation. Even with a good skeleton to start with, it’s a non-negligible amount of work to get that set up. I feel like the scenario in which this is useful is kinda narrow: when I have a very good idea of exactly what I want, but some step along the way is a hassle. General software engineering, like making a whole app, is far too open ended, and most of the sub-problems I encounter in software engineering seem either too open ended or too small to benefit from this approach.
That also doesn’t account for the speed of models. My experience is a ~20GB locally hosted model takes like 1-5 minutes to produce a good length response, and the few times I have used online models they are often slower. A few minutes per iteration, accounting for debugging when it goes off track, is not exactly fast or hassle free.
I just feel like the trade off where using AI vs doing it all myself is pretty limited in when AI offers an advantage.