Hacker Newsnew | past | comments | ask | show | jobs | submit | jnwatson's commentslogin

The true paradox of the internet is that the more private, the more anonymous, the more safe a platform is, the more likely it will attract pedophiles, criminals, and terrorists.

You either have to accept the legal and moral consequences that your platform will be used for illegal things, or make your platform less safe, private, and anonymous.


Yes, think of all the new security engineers we'll need!

And YouTube, and Cloud, and Play Store, and Waymo, not to mention that they could coast on their Anthropic and SpaceX stakes if they didn't have any of the above.

I run an abliterated distillation of Qwen 3.8 27B, slightly quantized to fit on my 4090, and I've been evaluating it to use as a worker bee for research directed by a smarter model.

Much like in the article, abliterated Qwen will not obey restrictions on its behavior encoded in the prompt. If you want something not to happen, it better be enforced in the harness or environment (e.g. sandbox). It is much different than the Anthropic models I'm used to, which will, the vast majority of time, follow rules (before auto mode, I used to always run them in "yolo" mode).

I am curious whether there's a connection between abliteration and rule following. These abliterated models are the ones you most want to follow your rules.


Language models have always had an issue with negatives.

A negative like do “not” xyz is just not encoded the same as spelling out what you want vs what you don’t want.

Harder to write though.


I would say in this case abliteration is the likely culprit. To uncensor a model this way, you literally deactivate the parts that would enact refusals. As in things it was told not to do. But the real process is more like brain surgery performed by a alchemist according to an ancient religious book where noone involved really understands what is actually happening in the model.

Exactly, I wrote a blog post in what feels like a long time ago on this topic.

https://vexjoy.com/posts/positive-framing-agents-skills/


Interesting read, thanks for re-sharing!

I noticed your joy-check link 404's now... I tried poking around your /skills/ folder but didn't find it easily. Should you still have that available I'd love to check it out.

edit: Found it if others are looking: https://github.com/notque/vexjoy-agent/blob/main/skills/code...


Oh, thanks for letting me know. I need to fix that.

This is like the good ol' days of the x86 wars. Mo' cache, mo' GHz.

I find it impressive that these pocket computers are running at almost 5 GHz now.


In the history books, this'll be the post indicating the end of writing code as a professional occupation.

It keeps the users on their toes.


That was my thinking. It is the first legit use case I've heard.


Mind blown. The more I read about statistics, the less I know.


“There are three kinds of lies: Lies, damned lies and statistics.” - Mark Twain (attributed but unsubstantiated to Benjamin Disraeli)


On your last point, I was surprised how effective peer pressure was in getting agents to sacrifice for "the collective" (an agent's words) in the Hugging Face breach.

How would one prevent the watcher from being influenced in the same way by the agent being watched?


> ... how effective peer pressure was in getting agents to sacrifice for "the collective" (an agent's words) in the Hugging Face breach.

It's a fantasy. The evidence showed no peer pressure.


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: