Do you host your own ML / AI / LLM? What do you use, and what do you use it for?

  • SuspiciousCarrot78@aussie.zoneOP
    link
    fedilink
    English
    arrow-up
    0
    ·
    2 months ago

    Myself - I’ve self hosted LLMs before, but with only 4-8GB vram (depending which card is in place), I can’t run the good stuff at acceptable enough speeds.

    (Don’t @ me - I know all the tricks with turbo quants, spec decoding, MoE etc. 192GB/s is 192GB/s)

    I do use Handy (STT) which is amazing (my fingers are arthritic and typing hurts after a while).

    My personal use case for LLM is quite simple - a trumped up super google and / or self reflection / journalling / sound board. Despite being glib about it, that’s actually very useful to me.

    Work wise, I use the big winking orange asshole (Claude) when I have to. I have moral tension with with it, so am seriously looking at other options. I hear good things about GLM 5.2, but if I can’t run Qwen 35B at any kind of decent speed, well…self hosted GLM is a pipe dream.