Hacker Newsnew | past | comments | ask | show | jobs | submit | csande17's commentslogin

Even if you take that website at face value, the ELO scores shown are relative to the other AI models tested, and not comparable to the ELO scores of humans who play against other humans.

I wonder why they didn’t throw a real chess engine in there for a baseline. There are engines where you can set the elo in the settings, so it should possible to see these LLMs relative to a human 1500 rather than just relative to each other.

> so it should possible to see these LLMs relative to a human 1500 rather than just relative to each other

As a 1500 elo human I can tell you that a 1500 elo chess engine doesn't play like anything like a 1500 elo human.


This is true, but I'm not sure it matters? I was poking around at the lichess database recently and those elo calibrated bots are remarkably well calibrated, their rating variance sticks out like a sore thumb compared to human players even at similar game volumes. So it should still be a decent predictor of how good a human at that level is, even if the playstyle seems alien.

I feel like every position is in the database so you could just lookup the most popular move for an arbitrary elo and that's the bot.

That would only work for the first few (from around 10 to 20 typically depending on how close people stick to opening book) moves.

Conservatively there are well over 10 to the 30 positions likely to show up in realistic games.

There are of the order of 10 to the 10 or so games recorded.

Thus well under one in a trillion positions are "known".


FWIW, I also got a bit of an LLM vibe skimming the changelog. In particular, the section about Linux sandboxing jumped out at me:

> Status in 7.0.0: kernels without Landlock continue working without Linux sandboxing in the less secure pre-6.0.0 configuration; brew doctor reports missing protection as an advisory.

> No replacement opt-out; unavailable Landlock remains advisory.

With that being said, I can kind of see why "eliminate all traces of AI writing styles" wouldn't be a goal of the review/editing process. I've been out of the Apple ecosystem for a while, but I imagine Homebrew has enthusiastically adopted LLMs for a lot of tasks. If the public communications reflect that, then people who don't like LLMs can bounce immediately, rather than getting invested in the project and feeling betrayed later.


If you say "AI tools", some vibe-coders try to equivocate between tools that generate the whole program for you and, like, an editor that uses an scoring algorithm to decide which method to show at the top of an autocomplete list. But if you say "vibe coding", some vibe-coders try to claim that what they're doing technically isn't vibe coding because they applied some non-zero amount of testing or review during the process.

If you ask "did you build this using Codex or Claude", it removes most of the wiggle room. And it's worded pretty neutrally, so it will often get vibe-coders who aren't technically using either of those tools to say something like, "Actually, I've built a custom harness for DeepSeek..." instead of an outright denial.


You're absolutely right! Let's delve deeper...

There were too many meetings, so they got rid of the people whose jobs were to create lots of meetings.


The devlog is the very first place changes get written about when they land in Zig's master branch (and sometimes even before!), so presumably zig.guide will mention this change once it actually makes it into a release.


Ideally, the documentation changes would have been merged together with the code changes.

In practice, Zig has very poor documentation overall.


The documentation is the most 0.x piece of Zig. It will come but it will take time.


Yeah, Android's sandbox deals with low-level file access and socket APIs and stuff. But Android also, intentionally, allows apps to expose their data and functionality to other apps via the higher-level "intent" system.

Apps get to choose what permissions are needed to access their intents, so under Android's security model, really it's Chrome's fault (or whatever browser the user has installed) for exposing an intent that allows apps that don't have the INTERNET permission to exfoltrate data.

Similarly, apps are also allowed to collude to share data with each other if they want; that's how stuff like Google Play Services works.


> Apps get to choose what permissions are needed to access their intents, so under Android's security model, really it's Chrome's fault (or whatever browser the user has installed) for exposing an intent that allows apps that don't have the INTERNET permission to exfoltrate data.

You might as well blame the phone for not being encased in concrete and thrown into a well. You might not want your text message application to have the internet permission, but you'd certainly want to be able to open links from it.


Unity includes its own video player implementation ( https://docs.unity3d.com/Manual/Video.html ), so presumably you'd mainly want to use VLC for wider range of supported codecs?


Basically. Unity video player is fine if you control the content. The other very reasonable option is to funnel shared content through a synced web browser interface. I've developed a few collaborative XR applications, and have used a blend of things to support sharing.


It’s not hard to use libavcodec to transcode, though.


https://docs.kernel.org/process/cve.html

> Note, due to the layer at which the Linux kernel is in a system, almost any bug might be exploitable to compromise the security of the kernel, but the possibility of exploitation is often not evident when the bug is fixed. Because of this, the CVE assignment team is overly cautious and assign CVE numbers to any bugfix that they identify. This explains the seemingly large number of CVEs that are issued by the Linux kernel team.

(And because this happens during the stable release process, there are a lot of 24-hour periods where they issue a ton of CVEs for all the minor bugs fixed in the release.)


"Anarchy" in Minecraft means no rules, and "you can't say bad words (or build things containing bad words)" is a rule.


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: