Hacker Newsnew | past | comments | ask | show | jobs | submit | sigmar's commentslogin

>Just remove hacking (bio-weapon, etc.) data from the training dataset and you're done.

Reasoning about how to write secure software uses the same knowledge as reasoning about how to break/hack it.


Just remove anything software-related from the training dataset.

Which also solves the alignment problem with those who do not enjoy seeing LLMs write software. But that brings us back to: Aligned to whom?


I don’t buy it.

Reasonable about building secure software can take the form “this memory access might be out of bounds — that MUST be fixed” or “this process has access to an inappropriate privilege — this is a serious weakness”.

Exploiting things and the capabilities that the labs call “cyber” are about the ability to (a) find the issues mentioned above and then (b) string issues together and avoid all the imperfect mitigations to actually compromise something. That latter part was IMO not actually necessary to train extensively, and I’d be quite happy to use a model that has no special skills in this regard but that would do (a) without complaining.


Lots of private benchmarks already exist, where you have to trust the tester (ex Artificial Analysis, Arc-agi).

>To understand the truth about the Hugging Face hack, you could do a lot worse than to listen to Ed Zitron and Cal Newport's recent podcast conversation

Lol, okay...

The crux of this piece is Doctorow saying that the hack was just a stochastic parrot repeating steps it has been trained on. Who cares how the LLM learned to hack things? Doesn't really change the facts of what happened. "Oh, it only made those paperclips because it saw instructions on making paper clips in the training data." These are some 2024 arguments...


I can't help but notice your counter arguments are "lol, okay" and "These are some 2024 arguments".

Not exactly convincing stuff.


Sure, if you remove almost all of my comment my argument disappears.

To spell things out for people that don't know about the topic: Zitron is neither an expert on the topic, nor a credible source of information: https://techreport.ngo/ai-ml/how-accurate-have-ed-zitron-s-a...

The paperclip maximizer is a thought experiment. If I tell an AI to start producing paperclips, it might start producing paperclips by doing unintended things. Technically, recycling the metal from all the world's bridges would assist in making more paperclips, but I never intended that. That's where my analogy to the post comes from- Openai intended the model to hack, but did not intend for it to hack HF. Do you think the fact 'recycling metal into paperclips' may have been in the training data is relevant to the thought experiment? Similarly here, it's orthogonal to the lessons from the HF hack and only brought up here seemingly to make it a fight over whether the models are really "autonomous"


> Who cares how the LLM learned to hack things? Doesn't really change the facts of what happened.

Seems like a pretty clear and succinct argument. If your only counter is to complain about tone, you’re losing.


That is all that Ed Zitron deserves lol. Completely unserious person. Really on the same intellectual level as the flat earthers at this point. And at both we may simply point and laugh.

It’s shared context for anyone familiar with the field.

Amusimgly, also how cults and fascism works.

They also breathe air. What’s your point? Don’t exercise reading comprehension because bad people do it too?

Yeah, basically any human group has some shared social context. Most cults and fascists probably eat together sometimes too, but that doesn't make it a red flag.

Ok; but the other side is equally delusional.

> "Oh we told the AI to use the tools, as well as to not use those tools. It chose to use the tools - we consider this cheating (for neabulous reasons), so lets get everybody in a panic about the morality and ethics, and how we can program those into the AI."

We know perfectly well how to constraint these programs. Attack isn't growing faster than defense. The people who believe in existential risk and want to teach AI's to be nice, as the last line of defense aren't helping at all. They're just jumping on the fearmongering bandwagon, perpetuating an "other consciousness" misunderstanding of the tool.

I've not seen LLMs display competence we should be fearful of the damage _it_ will do if left unchecked. All the damage will be done by ourselves to ourselves, regardless of the safeguards ideas being floated about.

My current belief is this whole HF media circus started with the simple human desire of OpenAI engineers to frame it such, that nobody would question their incompetence & liability & complicity.

Nobody is ever held responsible for out of control forces of natural powers after all.


What is a better/more faithful articulation of the incident in your mind?

In the podcast, Cal Newport repeats at least 20 times that agentic AI is nothing more than an LLM execution loop, as if this is some kind of trump card, taking for granted the usual stochastic parrot trope. At no point does he even consider the possibility that the LLM output itself is noteworthy. These people are permanently stuck in November 2022.

At this point, the "stochastic parrot" people are sounding like they need their prng seeded with less predictable numbers.

The output hasn't changed because the input hasn't changed. The shoe still fits.

So far everything an llm has ever done is still consistent with fitting bits of training data together.

It looks like people because it is replaying things done by people.

Similarly in the other direction, the fact that people can and often do mechanical things (make bad art, follow routines, etc) does not prove that people are no different than machines either.

Just because there are these two overlaps in both directions doesn't excuse getting them actually confused.

I don't think anyone lacking the perception to distinguish these things simply because there are overlaps and similar appearances is in a great position from which to be calling anyone else stochastic.


>or there is a more deliberative approach to assigning credit than who was "first" to solve some problem

it feels like an unintended consequence of the millennium prize is that people view the [last contributor to the solution] as the only one to make progress on the problem. I've never viewed Poincaré as solved by one person and the objective of the prize was to encourage more people to make attempts and contribute towards progress.

this issue is independent, but in these circumstances perhaps interweaved, with the 'ai is taking over math' concerns


Were the agents ever tasked with algorithm improvements? Post just says he didn't find any ("report essentially no algorithm advancements"). These LLMs are useful for optimization tasks where they can attempt a change and then measure performance boosts, so just wondering.

>We have a lot of physics based simulation tools, but they tend to focus on small subsets of the full design problem and they make limiting approximations

Do you think it is possible that better math will lead to better physics models?


It might but the math results from GenAI so far have been limited to finding counterexamples to known conjectures, not building new mathematics.

> results from GenAI so far have been limited to finding counterexamples

Not all.

Ehrhart’s volume conjecture

Quantum parallel repetition for general two-player quantum games

Erdős Problem #183 on multicolor Ramsey numbers

Erdős–Sárközy Problem #12(i)/(ii)

Erdős Problem #125

Log-concavity of codimension-3, type-2 pure O-sequences

Optimal O(1/t) last-iterate convergence for Anchored Gradient Descent-Ascent


True, these results are from last month. My info was a little out of date.

Prior updated :)


Yes, definitely! There’s a long history of this and I think there’s tons of opportunities for more. Both for improving the exactness/physical fidelity of models and for developing new approximate theories and simulation methods

>The speed with which we were able to produce this proof demonstrates that it is now possible to formalize large swaths of mathematics, which may both catch errors in the common body of mathematical proofs and reduce the burden of refereeing new work.

^ this section should have been in the first few paragraphs imho. Explaining why this is relevant shouldn't be so far down.


Forgive the authors of the article for assuming readers would complete it.

For any body of text (or in general, any exposition of any kind), the responsibility to explain the value of the article is very much in the author's side.

Explaining the value of what you are showing should always go towards the start. Else, why would anyone bother with the rest?


Feel very grateful I was never taught this... Would have missed out on quite a lot of good bodies of text in my life I think! Pushing through any initial friction or ignorance I might have as a reader, having the patience and charity to bear with an author until you get it, was instead what I was always taught.

Giving such a blanket "responsibility" to the author at all is just such a bummer! I say let them do whatever they want, there is always more than one way to express oneself. Someone who was never taught to write a clear thesis in the first paragraph for whatever reason doesn't inherently have less to say.


Unfortunately I don't see this particular view paying off in the age of AI, as many prove they have nothing at all to say but say it anyways. Which isn't to say people shouldn't write if they enjoy writing, but I for one will stay a discerning reader.

> Feel very grateful I was never taught this...

Never heard of Abstract section? First semester on a college or last year on high school.


Isn’t deciding what responsibilities there are in any text very much in the author’s side and not yours?

Haven’t you dramatically overstated your case? Many expositions do not contain an explanation of their value at all. Works of fiction are a good example, and there are many many others. Often it’s the responsibility of the readers & reviewers to decide on questions like value.


I think this is a Fermat joke :-)

Buzzard is writing for his blog audience - mostly mathematicians and not the casual visiting HN user.

Eh? The quote is from Anthropic, not Buzzard.

Have you written text for humans? You’re lucky if you can get people to read more than the title. You’re very lucky if they read past the first paragraph.

you would never assume this if you've spoken to any human being, ever

> reduce the burden of refereeing new work.

As a professional mathematician, I rarely need to worry about the correctness of a paper. The main difficulty of writing a review is instead understanding what the results of the paper mean in its context, how the results are presented, etc.


Isn't it the cost we care about, rather than the speed? All we know know is that a frontier AI lab was able to do it in 11 days, we have no idea how much compute they threw at it.

They said 6 billion tokens, which isn't as much as I thought it might be.

Am I doing my napkin math correct? The post says it's using a model comparable to Fable 5.1, which is $50 per million output tokens. So this is ~$300K? Surely an over-estimate due to caching.

Surely input tokens are also involved, and not necessarily only for the initial prompt if there are feedback loops or agent interactions.

Nah they should have released it in a 14-part tweet instead.

https://en.wikipedia.org/wiki/Attempted_assassination_of_Don...

what evidence is there that this registered republican was "far left"?


Well, there is apparently no evidence that he was "far right" either.

    So far, investigators haven't found any evidence on social media or other
    writings by Crooks that might help identify his motive for the attempted
    assassination, law enforcement officials say... And a review of public
    records suggests he may have had divergent political leanings, with Crooks
    registering to vote as a Republican but making a small donation to a
    Democratic-leaning group.


>We are currently working closely with our inference partners and open-source maintainers to align the technical details and ensure the model can be reliably deployed across the ecosystem. The full model weights will be released by July 27, 2026. Further details regarding the architecture, training, and evaluation will be released with the Kimi K3 technical report.

(translated by chrome)

11 days is a long time. It does not take that long to implement inference at providers. In my opinion, seems like they're being pre-emptively cautious about government intervention/review


Actually it does for a massive model, serving it correctly is not easy.

I believe Kimi also does some sort of Q&A and eval for day 0 partners, since early on a long of inference providers just weren’t running their models properly.


Eh, Minimax M2.7 also took a similar amount of time (actually longer) between availability and weights release.


Most of what an LLM does "could have" been done by a human if you throw enough human hours at it. But the reality in this circumstance is that a new tool helped find this leak. Saying this could have happened in a "non LLM world" is analogous to "someone else could have discovered special relativity, let's not mention Einstein"


This not only could have happened pre-llm, it did: https://krebsonsecurity.com/2022/02/report-missouri-governor...


My point is about the emphasis of Codex in the title. That emphasis makes more sense when Codex is credited with finding something that would have been difficult or impractical to discover without substantial human effort.


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: