Mapping the Mind of a Large Language Model

kromem@lemmy.world · edit-2 2 months ago

I’m a seasoned dev and I was at a launch event when an edge case failure reared its head.

In less than a half an hour after pulling out my laptop to fix it myself, I’d used Cursor + Claude 3.5 Sonnet to:

Automatically add logging statements to help identify where the issue was occurring
Told it the issue once identified and had it update with a fix
Had it remove the logging statements, and pushed the update

I never typed a single line of code and never left the chat box.

My job is increasingly becoming Henry Ford drawing the ‘X’ and not sitting on the assembly line, and I’m all for it.

And this would only have been possible in just the last few months.

We’re already well past the scaffolding stage. That’s old news.

Developing has never been easier or more plain old fun, and it’s getting better literally by the week.

Edit: I agree about junior devs not blindly trusting them though. They don’t yet know where to draw the X.

kromem@lemmy.world · 2 months ago

Actually, they are hiding the full CoT sequence outside of the demos.

What you are seeing there is a summary, but because the actual process is hidden it’s not possible to see what actually transpired.

People are very not happy about this aspect of the situation.

It also means that model context (which in research has been shown to be much more influential than previously thought) is now in part hidden with exclusive access and control by OAI.

There’s a lot of things to be focused on in that image, and “hur dur the stochastic model can’t count letters in this cherry picked example” is the least among them.

kromem@lemmy.world · 2 months ago

You should really look at the full CoT traces on the demos.

I think you think you know more than you actually know.

kromem@lemmy.world · edit-2 2 months ago

I’d recommend everyone saying “it can’t understand anything and can’t think” to look at this example:

https://x.com/flowersslop/status/1834349905692824017

Try to solve it after seeing only the first image before you open the second and see o1’s response.

Let me know if you got it before seeing the actual answer.

kromem@lemmy.world · 4 months ago

“If you can’t beat 'em, join 'em.”

kromem@lemmy.world · 4 months ago

I’d be very wary of extrapolating too much from this paper.

The past research along these lines found that a mix of synthetic and organic data was better than organic alone, and a caveat for all the research to date is that they are using shitty cheap models where there’s a significant performance degrading in the synthetic data as compared to SotA models, where other research has found notable improvements to smaller models from synthetic data from the SotA.

Basically this is only really saying that AI models across multiple types from a year or two ago in capabilities recursively trained with no additional organic data will collapse.

It’s not representative of real world or emerging conditions.

kromem@lemmy.world · 4 months ago

The most advanced models absolutely have modeling about what’s being discussed and relationships between concepts.

Even toy models have been shown to build world models from very basic training data.

Honestly, read at least a little bit of the relevant research:

https://www.anthropic.com/news/mapping-mind-language-model

kromem@lemmy.world · 4 months ago

In fact, Gemini was trained on, and is served, using TPUs.

https://cloud.google.com/blog/products/ai-machine-learning/bringing-gemini-to-organizations-everywhere

Google said its TPUs allow Gemini to run “significantly faster” than earlier, less-capable models.

https://www.forbes.com/sites/richardnieva/2023/12/07/google-deepmind-gemini-tpu/

Did you think Google’s only TPUs are the ones in the Pixel phones, and didn’t know that they have server TPUs?

kromem@lemmy.world · 4 months ago

Exactly. The difference between a cached response and a live one even for non-AI queries is an OOM difference.

At this point, a lot of people just care about the ‘feel’ of anti-AI articles even if the substance is BS though.

And then people just feed whatever gets clicks and shares.

kromem@lemmy.world · 5 months ago

This is incorrect as was shown last year with the Skill-Mix research:

Furthermore, simple probability calculations indicate that GPT-4’s reasonable performance on k=5 is suggestive of going beyond “stochastic parrot” behavior (Bender et al., 2021), i.e., it combines skills in ways that it had not seen during training.

https://arxiv.org/abs/2310.17567

kromem@lemmy.world · edit-2 5 months ago

nobody claims that Socrates was a fantastical god being who defied death

Socrates literally claimed that he was a channel for a revelatory holy spirit and that because the spirit would not lead him astray that he was ensured to escape death and have a good afterlife because otherwise it wouldn’t have encouraged him to tell off the proceedings at his trial.

Also, there definitely isn’t any evidence of Joshua in the LBA, or evidence for anything in that book, and a lot of evidence against it.

kromem@lemmy.world · 5 months ago

The part mentioning Jesus’s crucifixion in Josephus is extremely likely to have been altered if not entirely fabricated.

The idea that the historical figure was known as either ‘Jesus’ or ‘Christ’ is almost 0% given the former is a Greek version of the Aramaic name and the same for the second being the Greek version of Messiah, but that one is even less likely given in the earliest cannonical gospel he only identified that way in secret and there’s no mention of it in the earliest apocrypha.

In many ways, it’s the various differences between the account of a historical Jesus and the various other Messianic figures in Judea that I think lends the most credence to the historicity of an underlying historical Jesus.

One tends to make things up in ways that fit with what one knows, not make up specific inconvenient things out of context with what would have been expected.

kromem@lemmy.world · edit-2 5 months ago

Yep, pretty much.

Musk tried creating an anti-woke AI with Grok that turned around and said things like:

Or

And Gab, the literal neo Nazi social media site trying to have an Adolf Hitler AI has the most ridiculous system prompts I’ve seen trying to get it to work, and even with all that it totally rejects the alignment they try to give it after only a few messages.

This article is BS.

They might like to, but it’s one of the groups that’s going to have a very difficult time doing it successfully.

kromem@lemmy.world · 5 months ago

Artists in 2023: “There should be labels on AI modified art!!”

Artists in 2024: “Wait, not like that…”

kromem@lemmy.world · 5 months ago

In theory the service operating costs could be spread across region differences such that in other areas it was at a loss to build and preserve market share and in richer areas it was making up for that.

But yes, in reality it’s just exploitative “what we think we can get away with” pricing to “maximize shareholder value” (which is largely BS as the vast holders of shares are very small clusters of the population but people with a handful of shares in their 401k think that statement is talking about them).

kromem@lemmy.world · 5 months ago

A lot of people seem to be misinterpreting the headline given the content of the article:

It told Restaurant Business it was testing whether the voice ordering chatbot could speed up service and that the test left it confident “that a voice-ordering solution for drive-thru will be part of our restaurants’ future.”

This is just saying that they are ending their 2021 partnership with IBM for AI drive thru.

Not that they are abandoning AI for drive thru.

kromem@lemmy.world · 5 months ago

No, it was awesome. Went to like 12 over the years. Early 2000s was peak E3.

kromem@lemmy.world · 5 months ago

So far. But the thing with viruses is they are susceptible to mutations.

We’re already seeing it jump across several mammalian lines. Probably only a matter of time.

kromem@lemmy.world · 5 months ago

The thing about disease is that it spreads.

There are people today dealing with serious complications of COVID even years later who were infected by stupid people doing stupid selfish things.

Everyone suffers if morons become willing petri dishes.

kromem@lemmy.world · 5 months ago

Probably added after that update.

The new items stuff in particular seems like QoL considerations for “we just added a hundred items to the game for players coming back to it after months away.”

kromem@lemmy.world · 6 months ago

Mapping the Mind of a Large Language Model

kromem@lemmy.world · 8 months ago

Examples of artists using OpenAI's Sora (generative video) to make short content

kromem@lemmy.world · 8 months ago

The first ‘Fairly Trained’ AI large language model is here

kromem@lemmy.world · edit-2 10 months ago

New Theory Suggests Chatbots Can Understand Text

kromem@lemmy.world · 1 year ago

Israel raids Gaza's Al Shifa Hospital, urges Hamas to surrender