I’m curious what you are using. The free versions of chatgpt have been like that for me, but even Gemini flash with extended thinking, also free for a while longer, is giving me pretty reliable results as long as there training data out there to derive an answer from. The higher (paid) Claude models will one shot most coding tasks.
I’ve not yet fucked with Claude. I don’t want to pay for it, and I really don’t like the surveillance aspect of these centralized systems. Mostly I’m using Gemini, whatever DDG had in their search results, and local models I’ve been fiddling with (like Qwen3.8 right now).
All more or less garbage once I get into the details of anything on the edge of my expertise.
Are you using Gemini in flash extended thinking? (Not flash light) . I haven’t had many hallucinations other than cases where the training data it needs just doesn’t exist (cases where I can’t find the answers by googling either)
I’ve been using OpenCode with whateverthefuck free models they have listed on there and they all seem to do fine with agentic tasks like building me scripts or executables to make my work tasks easier.
I used Gemini at the start with “Frontier Knowledge” and it seemed to do worse than the ones listed on OpenCode, but maybe that’s because i could only do like three prompts a week since I refuse to pay into an AI.
end of the day, its just LLMs writing code for me, but I cannot see how this would be useful for a large scale code base, but also #NotAProgrammer.
Claude can one shot tasks until you get a larger system then it completely shits itself. These models are nothing more than autocomplete, and they can’t hold large systems in their heads. Anthropic literally tried to rewrite all of bun using Claude, they said they did it and yet it still hasn’t released six months later.
until you get to a larger system then it completely shits itself
This has gotten a lot better for me by having a “send out scouts” skill that has a lower tier model agent search through the codebase before it starts to plan. Has handled my companies giant monolith pretty well and even can handle cross repo features as well.
Yeah I certainly wouldn’t advocate building a whole business around code it wrote. But for small personal tasks it hasn’t let me down. Building custom server applications, desktop applications, Firefox add-ons, upgrading my homeassistant 10 versions over a couple weeks without letting anything break. These sort of things it handles pretty easily and are all things I wouldnt get done without it.
Gemini does a pretty good job of this because it doesn’t seem to have much built-in knowledge. Instead, it just searches the Internet on your behalf and returns summarized results with links to where it got that specific information.
I use it to search for scientific research all the time and the summaries often aren’t detailed enough so I actually click on those links. I’ve yet to encounter a situation where it fucked that up (invented links that don’t exist) but I have heard about it happening.
So far, the summaries have seemed to be pretty spot-on when it comes to biology papers 🤷
By writing instructions to insist that it double verifies every (non obvious) claim with a minimum of two independent sources. I also told mine to always assume that the initial prompt is missing crucial context, and to ask as many follow-up questions as necessary until it has enough information to provide the answer to the question I’m really asking. (For speed and efficiency you can even make it give you multiple choice options to click on.) Because sometimes the problem isn’t with the LLM, but with the user asking the wrong questions.
Using those two instructions alone, I’ve encountered considerably fewer hallucinations, and when I’m still not certain, I can simply click on the sources linked next to every single claim the AI makes.
Right, but that sounds like you implicitly trust the LLM and only verify statements when they seem off. So your intuition is the final arbiter if truth?
I’m curious what you are using. The free versions of chatgpt have been like that for me, but even Gemini flash with extended thinking, also free for a while longer, is giving me pretty reliable results as long as there training data out there to derive an answer from. The higher (paid) Claude models will one shot most coding tasks.
I’ve not yet fucked with Claude. I don’t want to pay for it, and I really don’t like the surveillance aspect of these centralized systems. Mostly I’m using Gemini, whatever DDG had in their search results, and local models I’ve been fiddling with (like Qwen3.8 right now).
All more or less garbage once I get into the details of anything on the edge of my expertise.
Are you using Gemini in flash extended thinking? (Not flash light) . I haven’t had many hallucinations other than cases where the training data it needs just doesn’t exist (cases where I can’t find the answers by googling either)
I got a free year of perplexity.ai which gives me limited access to claude sonnet and yeah, its honestly pretty good.
I’ve been using OpenCode with whateverthefuck free models they have listed on there and they all seem to do fine with agentic tasks like building me scripts or executables to make my work tasks easier.
I used Gemini at the start with “Frontier Knowledge” and it seemed to do worse than the ones listed on OpenCode, but maybe that’s because i could only do like three prompts a week since I refuse to pay into an AI.
end of the day, its just LLMs writing code for me, but I cannot see how this would be useful for a large scale code base, but also #NotAProgrammer.
Claude can one shot tasks until you get a larger system then it completely shits itself. These models are nothing more than autocomplete, and they can’t hold large systems in their heads. Anthropic literally tried to rewrite all of bun using Claude, they said they did it and yet it still hasn’t released six months later.
This has gotten a lot better for me by having a “send out scouts” skill that has a lower tier model agent search through the codebase before it starts to plan. Has handled my companies giant monolith pretty well and even can handle cross repo features as well.
Claude code has been using the new rust bun for ~5 months now and has been working fine, and bun 1.4 that released in August is using the rust rewrite
Yeah I certainly wouldn’t advocate building a whole business around code it wrote. But for small personal tasks it hasn’t let me down. Building custom server applications, desktop applications, Firefox add-ons, upgrading my homeassistant 10 versions over a couple weeks without letting anything break. These sort of things it handles pretty easily and are all things I wouldnt get done without it.
How do you verify that everything it tells you is correct?
You check the links/references it gives you.
Gemini does a pretty good job of this because it doesn’t seem to have much built-in knowledge. Instead, it just searches the Internet on your behalf and returns summarized results with links to where it got that specific information.
I use it to search for scientific research all the time and the summaries often aren’t detailed enough so I actually click on those links. I’ve yet to encounter a situation where it fucked that up (invented links that don’t exist) but I have heard about it happening.
So far, the summaries have seemed to be pretty spot-on when it comes to biology papers 🤷
By writing instructions to insist that it double verifies every (non obvious) claim with a minimum of two independent sources. I also told mine to always assume that the initial prompt is missing crucial context, and to ask as many follow-up questions as necessary until it has enough information to provide the answer to the question I’m really asking. (For speed and efficiency you can even make it give you multiple choice options to click on.) Because sometimes the problem isn’t with the LLM, but with the user asking the wrong questions.
Using those two instructions alone, I’ve encountered considerably fewer hallucinations, and when I’m still not certain, I can simply click on the sources linked next to every single claim the AI makes.
Right, but that sounds like you implicitly trust the LLM and only verify statements when they seem off. So your intuition is the final arbiter if truth?
It definitely has an element of garbage in, garbage out.