• PapaSkwat@lemmy.todayOP
    link
    fedilink
    arrow-up
    2
    arrow-down
    7
    ·
    edit-2
    10 days ago

    Good article on this, but I can not get this to work on my computer yet. I’ve been trying all morning. Okay, I finally got Qwen 3.8 27B working on my machine. I had to build a newer CUDA-enabled version of llama.cpp from source, adjust the GPU memory allocation, and limit the reasoning budget. It’s running now, and I’m still tweaking it for better speed. So far, so good.