• pigup@lemmy.world
    link
    fedilink
    arrow-up
    3
    arrow-down
    7
    ·
    12 days ago

    Call deepseek enjoy llm hallucinations and it jumping to conclusions and misconfiguring everything because it never bothers to check whether it’s just inventing bullshit, and there when it tries to solve the problem, it comes up with in correct theories and then further fucks everything up.

    Call GLM instead.

    • halcyoncmdr@piefed.social
      link
      fedilink
      English
      arrow-up
      16
      arrow-down
      3
      ·
      12 days ago

      You’re describing literally every LLM at some point. It’s why they can’t be trusted, and why the user needs to already know how to accomplish what they’re asking. The hallucination machine needs to only be used to try and make it faster, not relied on to blindly make decisions and changes without oversight.

      • pigup@lemmy.world
        link
        fedilink
        arrow-up
        1
        arrow-down
        4
        ·
        12 days ago

        Yea you better know what’s at stake when you /yolo

        Though progress is happening in the autonomy dept. Nearest models are more careful, and agent harnesses are improving. Maybe one day.

    • criticon@lemmy.ca
      link
      fedilink
      arrow-up
      11
      arrow-down
      1
      ·
      edit-2
      12 days ago

      Same as claude?

      We have to use Claude at my job, as in they check out token usage

      The other day I had to build a prototype using some development boards and decided to leave everything to claude. I gave it the datasets and told it what I needed to do and it created completely incorrect schematics, like hallucinating a LED on the board, adding a serial interface to some pins that are not related, confusing tx and rx on another set of pins, saying that it needed 12V when it is 3-5V, it even added some capacitors to some pins that weren’t even in use. I had to go thru maybe 10 iterations and burned a bunch of tokens, basically I had to do it manually

      The only thing it surprises me every time is the code, that usually runs with no errors on the first try. It’s just not very clean but it is ok for my needs

      • terranoid@lemmy.cafe
        link
        fedilink
        English
        arrow-up
        8
        arrow-down
        2
        ·
        12 days ago

        As a engineer, I think these tools really are useful for coding, not perfect but useful… And it would make sense, because they robbed the open source community blind to make them.

        There was a fuck ton of good training data of working code, and it is better at that than maybe more niche use cases where it’s sometimes surprisingly okay at it. But you give it a ton of good code and it’ll be decent at it… That shouldn’t be too surprising.

        I do see it make common mistakes like over engineering the shit out of a problem that could be simplified, but so do devs.

        I honestly think we got AGI and people refuse to admit it because it’s not the life changing technology that scifi made it out to be. I think there’s “meh” level AGI and logarithmic requirements to improve it, not exponential, and now they’re like “…it’s gonna cure cancer any day now”?

        Nah, it’s not. I am pretty sure the fact that Trump is calling it super intelligence is the news I needed to hear to know it’s never happening in our lifetime, and that it’s the lie they all want us to believe.