* Release new model that scores an arbitrary 100 on a benchmark
* Get everyone to talk about you as the first model to ever score 100 on the 100benchmark.
* Tune it down over time so that you end up only scoring 75 on the benchmark and people get used to it, gaslight them into thinking it never changed or that it's just a harness problem, they can't run the old version locally anyways to verify. This also cuts your costs in half. Your gross margin on API calls goes from 70% to 150%.
* Release new model that scores 120 on the benchmark and advertise it as 50% better than the current model, while it's only in practice a minor increment. Everyone praises it as the second coming of Jesus Christ.
* Get everyone to talk about you as the first model to ever score 120 on the 100benchmark.
Because frontier models are completely opaque. Doing a controlled test of "the same model" months apart is simply impossible if you don't work for that provider (and even then, may not be feasible). We know from external observation that model performance changes minute to minute, day to day and week to week for a variety of reasons: load balancing, inference hardware, and shared RAM pool to dozens of internal software settings each of which impact cost, latency, time-to-first-token, quality, veracity, tool use, etc.
Those software settings are being changed in real-time by an algorithm and those algorithms are being tweaked and A/B tested daily by the ~~performance~~ revenue optimization teams. On the hardware side the footprint a particular model is running on is materially changing, growing or being re-distributed across DCs ~weekly.
>gaslight them into thinking it never changed or that it's just a harness problem
Your benchmark didn't get 100 ? It's normal, it's not deterministic, and also your harness is wrong, and also you didn't do it when US users were offline, and also you got it wrong, and also we don't care about your results, the hivemind is speaking louder than you (also our bots are spamming more than you and drowning you out).
This very website has, at all times, a group of people saying "<Previous model> was never good enough for coding, but <current model> is the best thing and a game changer!" while the other goes "<current model> bad, <previous model> was better!". It's all vibes.
Apple product won't be the cheapest but it is a full package (CPU, RAM, VRAM/GPU, fast-storage, etc).
If you look at the pricing of a full (x86) AI workstation you'd need around the nvidia GPU, you'd approach $10k easily (and be using a ton more wattage too).
Always being five years behind everyone and doing its own boneheaded decisions that they inevitably have to walk back does tend to make Apple the Aztec in this case.
Mockery aside: it's all the same. They're all good enough for portraits where all you want is some depth of field & contrast, they're all good enough for landscapes where they capture colors pretty well, and they're all catastrophic past 5x zoom. You can take the top 10 of DXOMark (hell, the top 20) and you won't realistically find a difference unless you're pixel peeping. Which, we all know your phone pictures are just being sent compressed over Whatsapp, it doesn't matter. If you want camera quality pictures, get a camera. Even a 15 year old shitty Canon will take better pictures than the 18 Pro.
>Recently, a Vercel engineer called on harnesses to send the programming language the client prefers
oh would you look at that, Vercel suggesting to abuse how standard headers have been used for decades so it can send Accept-Language: rust because it's too lazy to ask for standardising an X-Prefers-Lang or anything else, and Shopify is here to shit on the internet too. Great.
Don’t worry, you don’t need to attack them – they do a great job of making themselves look ridiculous with their ignorant conversation. I would be so embarrassed if I had suggested that in public then subsequently discovered that the header doesn’t mean that at all.
People keep trying to systematize things for the tools they keep saying don't need systemization. No idea why you can't just put the documentation programming language context in the URL. "/docs/typescript/...", "/docs/python/..." etc.
I don't think the agents are struggling with the idea that different URLs have different responses, and that checking the sitemap is a good idea.
This is what happens when you let things get to your head. You work on a (highly inefficient, often broken) piece of extremely basic software by prompting a slot machine and hoping it works. Calm down and stop forcing your shitty framework on employees, you're the most replaceable cog of all.
Rossmann might be a useful tool in fighting against companies, that doesn't change the fact that he's a gigantic asshole who regularly harasses people he disagrees with. In the same way, if Stallman could stop eating his toenails in public and saying that Epstein's victims were totally willing, maybe he would avoid being on that list.
Stallman ate his toenail once. It must have been 20 years ago at this point, get over it. And he did not say anyone was totally willing. He said the victims probably presented themselves as willing. Learn to read and stop attributing words to someone incorrectly and completely out of context.
If a pimp threatens and intimidates a prostitute into having sex with someone and forces them to act like they are willing, are they not still being assaulted when the act of sex occurs? If an act of sex is assault, is the other person not assaulting?
Stallman comes across as discerning here because he's drawing a distinction between two things (sexually assaulting someone who is resisting vs sexually assaulting someone who is presenting as willing), but there is no real difference. The problem is not in his verbiage. The problem is that his underlying message is wrong.
He is trying to create a rationalization for why minsky's reputation should not be harmed. If Minsky flew to Epstein's private island and had sex with a very young girl of unknown age, his reputation should be harmed whether or not she "presented as willing."
Half of these people are trying to actively make life worse for you, or someone you know. We care about that stuff because otherwise people like DHH will want to get you shot if you're black, or people like Stallman will want to sexually harass you if you're a woman.
If you have the privilege of being able to forget about it: congratulations. If you're part of those potentially at risk: closing your eyes just makes their jobs easier.
Because DHH openly calls for murder of "immigrants" (read: non white people, even if they've been there for decades and been more productive members of society than DHH has ever been.), and it's not even a dogwhistle anymore.
>When wolves get out of control, you shoot them. When gypsies take over public spaces, you deport them. This isn't hard, it isn't cruel. It's the basic logic of self-preservation.
(in addition, the concept of Native <any european country> is nonsensical. Europe is at the confluence of migrations and has been for millenia. Attributing "whiteness" as a distinctive attribute of europeans is at best ignorant, and at worst another racist rally point. Go tell the greeks they're white, see how they like it.)
Americans have a problem understanding this because race and racism is so embedded in their culture. I just been reading and hearing about American *historians" who claim "Anglo-Saxon" is a racist term (in the context of discussing medieval history) because they anachronistically see it through the modern American lens of "whiteness" that is a result of American history and politics.
The future is an inappropriate OS with catastrophic battery management, zero software available, no sandboxing or proper security measures, and the oh so stable kernel ABI & Gnome APIs to develop for, yep yep yep.
The future is an inappropriate OS which spies on you and harvests your data, gets more and more locked down with every release restricting user freedom, and prioritises corporate interests over end user interest, yep yep yep.
I think many large cities, to keep this simple are failed because they dont serve the people but rather require the residents to be enslaved to keep it functioning.
We all live significantly longer, but at the same time we all are forced to give our precious time to the city and its ills, with commuting and traffic a simple example.
I hope I am fair in saying you probably consider Tokyo to be a success, and yet if you actually look how Tokyo residents live its actually a depressing place. Everybody rushes to school, work, unis and so on every day and come home tired. Rinse and repeat for the vast majority of days in a year.
Thats sad, goto YT and watch a few vids of youtubers catching trains or riding a bike or even just plain walking.
You will find videos filled with people rushing to their duties not out of choice but because the system forces them too. Very few scenes will have someone walking their dog, or riding a bike. I think its sad how few Japanese are able to have a day off, even the elderly work into their 80s and longer.
Thats a failure, too few people able to have a day off, thousands crossing roads, catching trains but nobody actually enjoying their hard earned time.
> I think many large cities, to keep this simple are failed because they dont serve the people but rather require the residents to be enslaved to keep it functioning.
People are usually not exactly forced at gunpoint to move to a big city. In fact, big cities usually get a lot of net-immigration.
> Thats sad, goto YT and watch a few vids of youtubers catching trains or riding a bike or even just plain walking.
I've been to Tokyo. It's ok. But then, I'm one of those weirdos who likes city living, and lived in London, Ankara, Sydney, Singapore etc.
> [...] but nobody actually enjoying their hard earned time.
Singapore isn't a city, it's a state. That happens to be the size of a city. You wouldn't say that New York is a failed state, because it doesn't make sense, and because the rest of the united states is able to feed it.
* Get everyone to talk about you as the first model to ever score 100 on the 100benchmark.
* Tune it down over time so that you end up only scoring 75 on the benchmark and people get used to it, gaslight them into thinking it never changed or that it's just a harness problem, they can't run the old version locally anyways to verify. This also cuts your costs in half. Your gross margin on API calls goes from 70% to 150%.
* Release new model that scores 120 on the benchmark and advertise it as 50% better than the current model, while it's only in practice a minor increment. Everyone praises it as the second coming of Jesus Christ.
* Get everyone to talk about you as the first model to ever score 120 on the 100benchmark.
Bis repetitae.
reply