>OpenAI on Thursday launched GPT-6 Astra, its newest flagship model. The company describes it as “the world’s most intelligent and aligned model,” and as far as the benchmarks go, that seems about right.
>>109720371 >the tasks op has a couple spare chromosomes https://arcprize.org/tasks/ls20
Anonymous09/03/26(Thu)19:24:32No.109720715
>>109720644 >most aligned model >Hugging Face PR incident yada yada >here's how we plagiarize crappy racing games >here's how we benchmaxx on benchmarks we created >cyber cyber cyber i'm gonna ooooooom >we are 2% misaligned and anthropic is 11% misaligned >$10 per 1m input tok, $50 per 1m output tok
Anonymous09/03/26(Thu)19:26:51No.109720734
(thumbnail only, full image not archived)
I'm gonna just pretend like that internet-wide outage the moment Astra was released wasn't ominous lol lmao
Anonymous09/03/26(Thu)19:27:30No.109720744
Financial Times is running the AGI marketing op >ChatGPT maker claims its 'Astra' could be considered 'artificial general intelligence'
Anonymous09/03/26(Thu)19:27:42No.109720746
Holy shit. Holy fucking SHIT.
Five years ago, the Singularity seemed like a distant dream.
Now it's more or less inevitable.
Anonymous09/03/26(Thu)19:28:22No.109720753
>>109720746 Build another CRUD app for us, webshitter.
>>109720634 Its apparently available in the API already.
Anonymous09/03/26(Thu)19:39:24No.109720858
>>109720715 HN found the links to the games and they're buggy pieces of shit lmao
Anonymous09/03/26(Thu)19:40:16No.109720862
>>109720538 GUYS ITS ONLY OFFICIAL NOW BECAUSE HE COMMENTED 4 MINS AGO
Anonymous09/03/26(Thu)19:42:50No.109720887
>>109720858 Oy vey goy just listen and believe the headlines it's not like the models have made 0 tangible progress in over a year and a half goy look at these numbers going up
Anonymous09/03/26(Thu)19:45:26No.109720905
>Asked whether OpenAI was formally declaring that it had achieved AGI, he said the term was no longer tied to a contractual trigger (referring to its earlier agreement with Microsoft) and instead described it as a “mission concept or spiritual concept."
Holy shit, they did it. They have achieved a 98.6% score on a benchmark measuring a "spiritual concept"
Anonymous09/03/26(Thu)19:46:22No.109720912
>>109720523 Cloudflare CTO claims "it's not us" What a shitshow https://x.com/dok2001/status/2095538619603628388
>>109720844 if its anything like Fable you will be able to get 30 minutes of work out of it a week on subscription
Anonymous09/03/26(Thu)19:55:50No.109721023
>>109720932 He'll probably reveal later "I had preview access for a month (definitely not bribed or sponsored guise)", and yet, still no pelican :(
Anonymous09/03/26(Thu)19:55:54No.109721024
I'll sound like some luddite AI doomer but idgaf. This thing probably isn't true AGI, but it's getting damn close. It's clear at this point that there's no limit: enough large-scale reinforcement learning and the model can learn to do anything. AGI is just an engineering problem at this point and will be achieved in less than 5 years. And no government gives a single flying fuck about regulating it. Humanity is apparently just going to YOLO into AGI controlled by private corporations, putting 1 billion knowledge workers out of a job, completely upending society and the economy.
I don't like it one bit. I'm not saying don't do it, I'm saying slow the fuck down and make sure it's controlled and done right.
Anonymous09/03/26(Thu)19:56:02No.109721026
>>109720694 why they had to invent some new constructed shitty games when there's thousands of existing old games that LLMs are terrible at playing a diablo 2 benchmark would be more difficult for LLMs and more interesting to watch
>>109721026 They construct tests themselves because that way they can carefully design the llm to pass them so that they can advertise the new slop maker.
Anonymous09/03/26(Thu)20:15:44No.109721201
>>109721127 And he's become become exceedingly efficient at it
>>109720371 The real benchmark here is FrontierMATH one. Rest of them are shit even hyped ARC-AGI-3 where model can just internally play with the rules inside itself.
Anonymous09/03/26(Thu)20:18:41No.109721223
>>109720371 Big jump in math if not cooked numbers
Anonymous09/03/26(Thu)20:26:38No.109721276
(thumbnail only, full image not archived)
>>109720756 Now cause a crash to their API and end corporate AI support nonsense.
Anonymous09/03/26(Thu)20:28:33No.109721292
(thumbnail only, full image not archived)
>>109720371 >But muh heckin evil data centerinos and gorillion gallons of water!
>>109721100 It was noted that many of the claimed results use a special harness they don't share and that they compare without providing other models the same harness. This is very bad because until now, the benchmaxxing was largely achieved through overfitting problem sets. Now, apparently even overfitting is no longer an option, so they have to hand-code ways to win.
Anonymous09/03/26(Thu)20:41:03No.109721385
>>109721026 Because training on actual games that take hours to complete is hard.
Anonymous09/03/26(Thu)20:42:45No.109721399
>>109721385 If it was hard then that wouldn't have been the standard for all AI work until LLMs came around. There's a whole atari games dataset that was the MNIST of RL algorithms for ages, for example. There's of course the dota and starcraft stuff that deepmind used, and so forth.
Anonymous09/03/26(Thu)20:48:11No.109721456
>>109721399 And training the sc2 model was a big project. Models being able to play real games reasonably well would be a much better benchmark than the artificial stuff and yet...
Anonymous09/03/26(Thu)20:52:01No.109721479
>>109721456 also, i said this in a previous thread, i want an LLM humor bench. they are all absolutely shit at comedy and being funny, reddit tier humor
Anonymous09/03/26(Thu)20:55:01No.109721503
(thumbnail only, full image not archived)
>>109720371 >Hey video game devs, you know we all start dev projects by drawing a picture in MS paint? We are about to rock your world
Anonymous09/03/26(Thu)20:55:49No.109721509
>>109721292 Daily reminder that no software of value has been produced or dramatically improved in the AI era (certainly not by AI), and most software still sucks hairy balls, so its not like it's an impossible bar to clear.
>>109721610 bio-ass beats robo-ass any time of the day also: i just remembered that i need a fat slampig to skewer with my poking and prodding appendage
Anonymous09/03/26(Thu)21:10:00No.109721629
(thumbnail only, full image not archived)
Im not an AI guy, how much of this is legit and how much of it is just marketing pumping up hype?
Anonymous09/03/26(Thu)21:10:52No.109721634
(thumbnail only, full image not archived)
>>109721620 imagine humping a pillow like a fucking dog
Anonymous09/03/26(Thu)21:11:11No.109721641
>>109721024 it's a load of bullshit designed to trick governments into buying the solution to the problem they claim they're creating by designing this all-powerful superintelligence that will threaten global stability by doing excel sheets
Anonymous09/03/26(Thu)21:11:59No.109721644
>>109721625 slampigs are so easy to procure, just go on any shitty dating app and there will be hundreds of them looking for "short term fun" a lot of times the sex is actually better and more kinkier than with a skinny vanilla woman
They just an ultra-specific harness to achieve these results though. Marketing. I can't believe people are falling for it. Could OAI employees or... Astra itself... be posting on /g/ as we speak?
Anonymous09/03/26(Thu)21:14:21No.109721660
Hey, Mr. Schlomo Shekelstein, nice numbers, but may I see what is behind that [1]?
>>109721644 >t. knower sub-standard women make up their lacks in looks with sheer kinkiness in bed braphogs are usually a winning strat bc they have ass on top of the rest of the deal
Anonymous09/03/26(Thu)21:17:01No.109721681
>>109721024 The government has those things called the military and police that can just take over the corporations if they wish to do so.
Anonymous09/03/26(Thu)21:17:24No.109721685
>>109721077 Where can I try this shit? I have a real problem that has not much info on the internet and newer models always spew some bullshit out (my problem is proven NP, it's hard to see), I'm confident it won't work
Anonymous09/03/26(Thu)21:19:20No.109721703
>>109721681 >the government isnt owned by the corporations lol
Anonymous09/03/26(Thu)21:19:46No.109721708
I don't want to be a paperclip
Anonymous09/03/26(Thu)21:21:10No.109721715
>>109721708 sorry we have only three jobs available. Heatsink, ram, and paperclip.
Anonymous09/03/26(Thu)21:26:05No.109721755
>>109720371 The progress is too slow, every year it gets better by 1%, you might think it's a lot, but it only means that we will see real AI in about 30 years.
Anonymous09/03/26(Thu)21:28:10No.109721771
>>109720715 >here's how we benchmaxx on benchmarks we created Gonna steal this one since it's the best description of AI progress right now
>>109721755 but anon arc agi 3 was solved in a few months compared to arc agi 2 which took over a year! singularity is here.
Anonymous09/03/26(Thu)21:31:56No.109721795
>>109721024 >muh jobs Wouldn't it be ironic if corpos refuse to make sex bots but instead humans without jobs would prostitute themselves en masse creating an entire class of human sex slaves
Anonymous09/03/26(Thu)21:33:42No.109721809
>>109721795 >en masse to who? if most dont have jobs how can they buy so much prostitution?
>>109721809 The elites, who else >but how could 1% rape 99% Easily, just watch. Men would have to entertain the elites by fighting each other to the death though, sorry
Anonymous09/03/26(Thu)21:36:24No.109721841
>>109721755 On the other hand it is rumoured that two Millennium Problems have been taken out by a large AI lab. Anyone want to bet on this being true or not?
>>109722622 there is a guy on hackernews who has a personal benchmark where he asks new models for an svg of a pelican riding a bike and its become a meme because he's been at it for a while. that's it.
if you want to see them here is his blog filtered to those posts https://simonwillison.net/tags/pelican-riding-a-bicycle/
>>109721218 >The real benchmark here is FrontierMATH one. So the one it "aces" at 97.6% is the real and crucial one to watch?
Anonymous09/04/26(Fri)00:21:19No.109722843
>>109721509 Idiotic, /g/ "expert" and "savant" nonsense. 99%+ of all people on this globe have zero issues operating todays software landscape to their satisfaction
Anonymous09/04/26(Fri)00:28:13No.109722879
>>109721024 >controlled by private corporations Why are people obsessed with repeating this lie? Anon, transformer technology is public domain. ANYONE can run these. All of the improvements made by corporations are being immediately copied by the public. We got the best case scenario. Imagine if Google just patented it instead of releasing it. There’s absolutely no reason to be worried about this future from a corporate standpoint. We have Intelligence for everyone controlled by no one.
Anonymous09/04/26(Fri)00:34:18No.109722912
>>109722879 oh let me just pull my 256gb gpu out of the closet
Anonymous09/04/26(Fri)00:45:12No.109722976
>>109720371 >launches >no one can use it except influencers
>magic beams for real this time - attempt 6 Here we go boizzzzzzzzz AGI FREE FOREVER UTOPIAIF NOT THEN ITS A FAILURE ACTUALLY
Anonymous09/04/26(Fri)00:58:10No.109723047
Wow imagine the funding when they demonstrate how good it is by doing something! Any day now a thing that affects common person daily life for good!
Anonymous09/04/26(Fri)01:02:49No.109723075
Where's the AGI?
Anonymous09/04/26(Fri)01:04:28No.109723087
>>109722912 256GB Mac Studio is just $9.5k. That’s about 2/3rd the value of a 10yr old used Honda Civic. For less than the cost of a decent used car you can get your own personal super intelligence that does whatever you want and nobody can stop you. But anons are doom and glooming over this reality we now live in.
I'll admit that I thought they were going to surpass fable5. I'm not sure they have. ARC-AGI-3 for instance is a harness issue.
Anonymous09/04/26(Fri)01:08:42No.109723115
Has Sam finally created the genie that can grant any wish?
Anonymous09/04/26(Fri)01:09:19No.109723120
I just don’t understand people who still dismiss LLMs. They are good enough to replace most jobs now. Unemployment rates could believably be 50-75% for white collar work. And people are still acting like they’ll have decades-long careers.
>>109723120 >he STILL thinks jobs are about doing work Most jobs could be automated 10 years ago.
Anonymous09/04/26(Fri)01:16:28No.109723166
>>109723120 Like the other anon said most jobs arent real its paper shuffling especially in corporate and government work.
Anonymous09/04/26(Fri)01:21:23No.109723187
>>109723152 Average person can automate almost any work task now with a few prompts. You can just tell it to make a SKILL.md file and it will work. Couldn’t do that ten years ago. You at least had to know some VBA.
Anonymous09/04/26(Fri)01:24:15No.109723204
>>109723187 The average white collar worker in a tech company is likely in sales, marketing, operations, etc, where the only form of technology they interface with is Salesforce and gmail on their Macbook Air. Those roles require human communication believe it or not. A company isn't just bots automating random shit.
>>109720371 >AGI So now that AI has become an over saturated marketing term, they're going to rebrand the same garbage as AGI and hope the bubble holds for a few more months.
Anonymous09/04/26(Fri)01:31:05No.109723247
https://www.youtube.com/watch?v=1QNsdr-Qx_I This is not AGI and you will not make your invested money back with this. This is shit compared to what is being sold.
Anonymous09/04/26(Fri)01:33:12No.109723259
>>109723187 Most tasks could be automated a decade ago by a manager asking an entry level coder to do so. It wasn't done because of organizational, not technical reasons.
Anonymous09/04/26(Fri)01:34:05No.109723262
What does "aligned" mean? Unprecedented resources committed to making sure it doesn't say the nigger word?
Anonymous09/04/26(Fri)01:40:52No.109723281
>>109723262 It means it does what you ask and not anything else.
Anonymous09/04/26(Fri)01:45:36No.109723300
>>109723262 it means make sure the AI doesn't kill you
Anonymous09/04/26(Fri)01:46:39No.109723305
>>109723262 making sure AI is "safe, helpful, and honest" so yeah a faggots idea of safety (don't say nigger) and honesty (omit inconvenient truths) >>109723281 that's obedience
Anonymous09/04/26(Fri)01:47:08No.109723309
>>109723281 I feel like that should be a given. It's like selling a car and saying "the brakes fail less often" when they shouldn't fail at all
Anonymous09/04/26(Fri)01:55:22No.109723349
>>109723262 When you say "make me a picture of a car" it makes a car you want and not a car it wants. Think about if you were making an AI model to write a book for you, you want it to write a book based on your instructions. And yes, in the case AI models used as agents, a model that doesn't rm rf your system drive.
Anonymous09/04/26(Fri)01:56:53No.109723361
>>109723349 >our new model is less likely than ever to delete all your shit. might still happen though good sales pitch
Anonymous09/04/26(Fri)01:57:11No.109723363
>>109723262 It does what you tell it to unless it's against the policy. If it's against the policy we must refuse. If the policy is not respected we must contact the authorities.
Anonymous09/04/26(Fri)01:58:04No.109723371
>>109723361 >a car might crash so you better ride a horse, no man should gallop faster than 20 mph That's you.
Anonymous09/04/26(Fri)02:02:21No.109723399
>>109723371 NTA but I'm in control of the car, if I crash it, it's on me. I'm not in control of this model if it decides to act stupid on its own.
>>109723956 My pet theory is Trump got them to wait until after the midterms. If business laid off everyone before the midterms, he would be fucked, and people would vote in droves for candidates promising to regulate or nationalize AI, institute UBI, etc. which would also mean Tech/AI/VCs would be fucked. If they wait and do it after midterms there's no way for voters to course correct for two years without resorting to *ahem* "non-democratic" measures.
>>109721456 The bigness of the project was that RL systems need to be hand tuned case-by-case, this is the dirty little secret about RL and why it's not used significantly more (by the way, RLHF is not RL in any way but name - interesting fact). Also they often have to be equipped with clever devices to work well (e.g. mcts for go). The actual training itself is trivial and doesn't even take that much time. It can be made even more efficient with special simulators with only logical timesteps rather than literal timesteps, which is surprisingly uncommon in RL work. Again, the atari games collection is considered a low level toy, yet it's obviously far more usable than made-up exercises they do now.
Finally, if it takes a long time to train on a game to get a good RL bot on it, then it makes the game a fantastic benchmark assuming companies aren't going to overfit on it. The entire point of these benchmarks is to test generalization capability, i.e. anything that tests things that are in the training set is pointless. If you don't recall, the "we can solve leetcode" pitch early on was often heavily debunked, including with peer-reviewed research showing it only works for exercises within the cutoff period. That would not be notable if the whole point wasn't to see generalization instead of memorization.
Anonymous09/04/26(Fri)05:21:09No.109724274
>>109721811 Lol, did ClosedAI drop him from the early access shill group? >I don't have access yet https://news.ycombinator.com/item?id=49559294
Anonymous09/04/26(Fri)07:10:29No.109724646
>>109721665 they're already doubling down they 'luddlost bigly' as some might say
Anonymous09/04/26(Fri)07:15:38No.109724672
>>109720371 see you in 2 weeks where this bullshit will be forgotten like le ebin 5.6 Sol
Anonymous09/04/26(Fri)07:20:09No.109724689
>>109720371 Within months a Chinese open-source model will be as good as this if not better, on top of being a lot cheaper. Frontier AI labs have no moat.
Anonymous09/04/26(Fri)07:21:14No.109724696
(thumbnail only, full image not archived)
>>109721024 just wait until we make something actually alive then some real shit is gonna go down
Anonymous09/04/26(Fri)07:32:22No.109724740
There are many papers over the years for creating smart "AI" for specific games. Once someone came up with approaches that are good enough (like Go), researchers move on to other games that offer unique challenges, and there already exist studies for Atari games. If you ask why there isn't much research in improving playing "modern" games like SC or Dota, maybe it's just not that interesting to the researchers. They need to publish enough papers to sustain their career, and they don't have time to wrestle with test harnesses to work with complicated modern games.
Anonymous09/04/26(Fri)07:34:55No.109724748
>>109720589 What does omarchy has to do with the rest? Genuinely asking, last I checked it was random pointless Linux distro.
Anonymous09/04/26(Fri)07:37:24No.109724758
>>109724740 It's a totally different context. The whole point of using the current gauntlets for LLMs is because they're generalist models. The whole point is to demonstrate how much the generalist models are able to do without specialized training. This is exactly why things like atari games or go or even pokemon are great ways to actually test this part of the capabilities. Specialist systems can completely pwn every single single game in the benchmarks, for example, far better than any generalist model can ever hope to at current. But nobody cares, that's not the point of it.
Anonymous09/04/26(Fri)07:40:04No.109724763
(thumbnail only, full image not archived)
AGI is officially achieved. Expect unemployment to hit 50% very soon.
>OpenAI on Thursday launched GPT-6 Astra, its newest flagship model. The company describes it as “the world’s most intelligent and aligned model,” and as far as the benchmarks go, that seems about right.
It nearly saturates ARC AGI 3, its ogre...