Hacker Newsnew | past | comments | ask | show | jobs | submit | gymbeaux's commentslogin

My understanding is both USB 4 and TB use PCIe tunneling. Surely at this price point we’re looking at a USB 4 device. Thunderbolt has historically cost more due to licensing, which is now free I believe, but there’s still specialized hardware and requirements manufacturers need to adhere to.


It seems to support USB4 with what they call "stock firmware". They haven't clarified what that means, but my assumption has been that it's a binary blob made by the manufacturer of their microcontroller that doesn't support all of their fancy debugging features (thus actually turning my into an Aliexpress special eGPU dock with a high quality buck converter).

It’s an incredibly dumb implementation and it feels like yet another case of “AI startup gettin’ while the gettin’s good”. It’s the dotcom bubble all over again. Do anything with AI and people will buy what you’re selling.

The RX 9060 has a TDP of 160W, but that’s the high end of what you can continuously draw from the 12v DC plug in a car, so it looks like they cap it at 100W. It has to be on the floor or under the passenger seat because it has to plug into that 12v outlet that’s usually in the front.

Things will get caught by the fan blades, and passengers will inadvertently kick it. Hopefully suddenly losing connection to the GPU won’t cause the car to hit a pedestrian or anything like that.


You forgot to ask "does it work". If the answer is yes, then what is the complaint? They can let other people 3d print cases and make wiring harnesses.


wouldn't you also need to worry about the voltage spikes when you're running on battery vs starting the car if you're running an ICE? and would you need to unplug it if you, say, need a jump? feels like you'd also need to install a whole UPS on the other side too if you care about the hardware longevity (not factoring in the sheer amount of dirt/leaves/etc that'll just get caught up in the card itself)


What would happen to Nvidia, Anthropic, OpenAI, if tomorrow someone released an open weights model on HuggingFace that matched performance and accuracy of Opus 5 running locally on an RTX 5070? That won’t happen tomorrow, but it will likely happen someday… what’s the plan beyond “don’t be the one holding the bags?”


If I could have shown up somewhere in 2022 with a Mac Studio M1 Max w/ 64GB of RAM running Qwen-3.6-27B or 35B-A3B, I would have pretty much been a demigod - to a degree far more impressive than being able to run Opus 5 locally today.

So yes, I think your scenario is likely to eventually happen, but there will be a much more powerful, capable frontier model then.


There’s no reason to assume frontier-level intelligence eventually collapses all the way onto a midrange consumer GPU. In fact, there are quite a few reasons not to assume that (information-theoretic constraints, etc).


There's no information theoretic constraint we know of that prevents this. You will almost surely win a Turing award if you can prove this.

It's almost a given that whatever is frontier intelligence today will run on a potato in a few years.


Kinda silly to follow your “prove it” challenge with an absurd claim you most certainly cannot prove, much less support with evidence.


It was not a "prove it" challenge.

I'm pointing out that there's no known information theoretic constraint about the impossibility of frontier AI models being improved to fit/run on a small GPU.

Please do not make up plausible sounding science facts.


Please do not assert I am making a claim I’m not making. Information-theoretic constraints exist. My comment does not require some specific, hard constraint to have been clearly defined, for my point to be valid.

If I were to say you could put a motorcycle in my car’s trunk, it would be perfect valid for me to say there are space constraints that make your idea unlikely. The same is true in this discussion, even though I have not computed the exact dimensions of the motorcycle and my car’s trunk.


Your claim was about frontier intelligence and midrange consumer GPUs. That's a pretty specific constraint.

Sure, there could be some point between a midrange consumer GPU and a pocket calculator where you can't fit enough 'intelligence'. But we really have no idea if the constraint is information theoretic or something completely different. Demonstrating that is the hard part, not finding the exact number of bits.

Talking about motorcycles in car trunks is just lazy false analogy here.


Then I give up. Best of luck.


I will not claim a 5070, but there is already evidence in nature that you can get very good general intelligence with an order of magnitude less wattage.

There are constraints of course- training takes way longer.


I wonder if we'll eventually find that Darwin style evolution gets us close to the global optima of intelligence given constraints like size and energy.

We don't really have the tools to reason about this stuff yet. Exciting times.


But, it could happen for a coding-focused model, or an accounting-focused model, etc. most tasks only need a subset of the total model to be done effectively.


Could you not say the exact same of image gen models? For those that haven't kept up with that domain, you can now efficiently run high quality image gen models on any plain old video card, with phenomenal results.


Core reasoning model with plugins for specialized tasks like "Pip install" developed using the new science of AI neurosurgery.


We could very well reach a point where models don't get better anymore, or where consumer models are good enough for 95% of the use-cases.


Workloads will inflate just as they have been. Remember when llm assisted development used to be good only for a function, then a whole file, then a handful of files, then a code base, then a full stack, etc etc etc.

People will claim to have “enough” even though they already have the equivalent of last years capabilities locally.


Those companies will be quick to copy the tech, inference cost would plummet and there is a greater chance that these companies could make it to solvency. At least in the short term. Long term it might not be so great as consume hardware catches up.


Nothing would really change IMO? 99% of users don't have anything like a RTX5070 (mobile especially).

Even if it did, it still doesn't make much economic sense running a model locally vs on a datacentre.

For example, I managed to just about squeeze a Q2 quant of Qwen 3.7 27b on my 9070XT. I get around 60tps decode (slightly faster prefill). _but_ it uses 300W of power to do so. At UK electricity rates of 30c/kWh this works out at something like 42c/MTok. I can get far far better models on openrouter cheaper than that, plus I'm not horrendously constrained on context length.


I dunno a lot of things said about AI economics sound like an IBM executive making reassuring statements about their terminal/mainframe business before the personal computer took off.

Like even if you run it in a datacenter in this scenario, you could do it on a cheap GPU instance in Azure, you still wouldnt need OpenAI or Anthropic specific clouds.

>uses 300W of power to do so.

There are plenty of people with phat electricity pipes in their on prem server rooms that have been vacated for cloud. Companies who want the benefits of AI but dont want the risk of sending their data to foreign API endpoints.


The analogy with IBM mainframe completely ignores Murphy's law which came up and lead to the small and fast chips we have today.

But Murphy's law is dead. No future chip will leapfrog easily current chips because we have reached hard phyical limits in chip density and downsizing. Huang's law by Jensen Huang focuses on something else and that is token performance per Watt at scale.

Blackwell needs double TDP than Hopper and Rubin again needs almost double TDP on a rack but in the end Rubin will be like 100x token performance per watt on a scaled data center. This means you have more energy need but you get multiples of token performance because you start scaling in the data center.

The local chip will never be able to keep up with the data center scaling economics. This is why everyone is so crazy about building data centers because they can see the economocs behind it.

What people don't seem to understand if tokens become more available and cheaper then not only more people can use them but a single person can use more as well. Why should you be limited to 1 AI agent? Why can't have you have multiple agents running on multiple devices daily for you?

This is why demand will grow exponentially with the growth of token economics. We have seen it for the last few years and much more is yet to come.


Do... do you mean Moores Law?

>The local chip will never be able to keep up with the data center scaling economics.

Assumes the software has been completely solved.

>Why should you be limited to 1 AI agent? Why can't have you have multiple agents running on multiple devices daily for you?

At some point we cap out the bandwidth of the human to keep up with their mistakes.


You're suggesting that if a very good and cheap AI model came out tomorrow everyone would rush out to rent Azure instances to run batch size 1 inference on their model?


I am suggesting that Azure and AWS would change course and push corporate customers towards more expensive, but more private options.


It seems very very unlikely that an Opus 5 matching local model that runs on a 5070 will be released within the next 5 years (I don't want to say "ever").

If it does happen then NVidia will sell a lot of 5070s though!


"eventually" is actually a function of frontier model capabilities. You only get Qwen6-27B when you have Opus 7 producing extremely high quality tokens for them to train on. So the market for local models is always significantly behind the frontier, by definition.


If you could run Opus 5 on a 5070 then the labs must have achieved RSI at that point


probably not all that much... the market would dip, just like every time a new open weights model gets announced. but hundreds of millions of people aren't going to immediately self-hosting their own models.

the biggest winner in that scenario would be ai providers, who suddenly have a capable model that they can serve much more efficiently. and the incumbents have a whole lot of compute. wouldn't anthropic and openAI just start offering that open weights model at prices that nobody else could compete with?


They could but then their valuation is no longer justifiable, which breaks a lot of things downstream (loans being the biggie). They'd rather lose money than start making money in a non defensible way.


I think this might be the core signal that it’s a bubble.


> on an RTX 5070

RTX 5070 prices go up ~N times. Nvidia makes more money because it's easier to make these things than it's to make a GB300.


inference is the cheap part; training is expensive. what compute infrastructure would train this mythical magic model?


Exactly. If OpenAI and Anthropic didn't have to train new models, they'd (probably) be instantly profitable and with good margins.


Inference time scaling means whoever had the most compute has the highest intelligence model.


I guess I'd like to understand the technical reasoning on how you think an how an Opus 5 could over time fit on an RTX 5070.


We’re in the phase of society where we will just build something like this ourselves using an LLM and it will be built exactly how we want it and we don’t have to worry about malicious code or forking it and hoping the owner will merge our PR… it’s sad but that’s where we are for anything the size of a menu bar that pings a URL every minute.


yeah totally! some version of "summarize what this tool does at a high level and then build my own but don't use any of the code" perhaps.


But they’re slower (and benchmark worse) than Opus, GPT, et al. Why?


Personally I don't care about common benchmarks as I don't find actual coding agent performance correlates strongly with them. One reason to use open-weight models is that they don't hide the reasoning, so you can do very aggressive context management in your harness to use significantly less tokens. Smaller prompts are faster since KV is N^2 plus it can be dramatically cheaper (if you balance your aggressive context management with maintaining the prefix cache as much as possible). Even paying for API prices directly and using agents as much as I want I spend less per month than the $200 I spent on a Claude MAX sub when I had it.


Kimi and DeepSeek are impressive for their size but still run slow on any hardware you or I would have. If you’re proposing we run those via a cloud service, I don’t see a reason to do that when it’s an inferior model and I still have to pay per token for it.


OpenCode Go provides a lot of usage for these models for the paltry sum of $10/month. Z.ai's coding plan provides a single-digit multiple of Claude Code's usage for a similar price and performance level. Kimi and DeepSeek models are hundreds of billions of parameters (or, in K3's case, >1T). Many of these models have Opus-level benchmarks and, as I pointed out previously, often practically outperform Anthropic models because they're more consistent.


You should run the maths once. Those tokens cost you much more than the hardware would. But yeah, CAPEX vs OPEX something something.


When I see new software now, my first thought is an LLM wrote it and it makes me not want to use it. I also assume there are bugs and it’s not particularly feature-rich.


It is becoming very hard not to use an LLM. I have been working on a project all this year. I started doing everything by hand, and I am using LLMs a little bit more every day. I suppose we want to use projects that were crafted with care, whether or not LLMs were used. But it is not easy to figure that out. It reminds me. 30 years ago, before code formatters were popular, you could tell good code just by looking at how carefully formatted it was.


> It is becoming very hard not to use an LLM.

How? It's not like it accidentally installs itself onto your machine and runs itself, or injects itself into your projects without you knowing.


and you think Bloomberg engineers dont use AI heavily?


my first thoughts too.


same and for some reason people are downvoting. why would we use vibecoded stuff with sensitive information?


What’s sensitive about displaying public information in pretty graphs?


LLMs are incredibly inefficient and can never be “100% accurate” so we have effectively seen what they can and can’t do already. If LLMs haven’t taken your job by now, they aren’t taking your job.


Everyone who could be considered a role model in my life has been “laid off” (fired) at least once. These are successful people. They’re all millionaires (not just the retired ones). We need to get over the fear and stigma of getting fired. It means nothing except your boss didn’t like you.


I think you are focusing on the survivorship bias in role models, as those successful people in your life who you view as role models are more visible to you (the younger person) because they succeeded.

You wouldn't pick someone who ended up a Uber driver, because they dont look to you as someone from your field to model your future on.

Ultimately, getting fired can drastically change your life. Either for the betterment or detriment to their lives.


I look at all the “role model” esque people in my life - they’ve all been fired from a job at least once. There are also plenty of people I know who I wouldn’t consider role model esque who are simply lousy employees (and probably get fired many more times than the “role model” types). I say role model types simply to mean “people who you would be surprised to learn were fired from their job”.


> We need to get over the fear [..] of getting fired.

You're speaking from a position of privilege.

Sure, a considerable number of people (If not a majority) here on Hacker News are decent earners and likely have enough savings to ride out a job loss for at least a few months.

But I recall my days while I was working a fast food job. If I lost my job, unless I found a new one within two weeks, I'd be homeless. When your paychecks are just barely covering the cost of living and saving is impossible, then getting fired is a genuine fear. It means losing the roof over your head and the food from the table.


Stigma meaning others judging you for being fired. High level “no one” wants to hire someone who was fired from their last job. I’m just saying in my experience everyone gets fired sooner or later, regardless of their abilities, job performance, social skills, et. Al.


>We need to get over the fear and stigma of getting fired. It means nothing except your boss didn’t like you.

It also means (in the US) you lose access to your healthcare.


> " It means nothing except your boss didn’t like you."

And in some cases, especially in our industry, that the funding ran out.


I never said anything about a role model, I said people having to do multiple roles in one. But I kind of agree with you, if your boss doesn't like you, you can literally build rockets to the moon and get laid off :^)

These days it is much harsher, if you are not socially accepted or liked, you are out of luck, your skill doesn't matter in their eyes, until... it does again


But what if all your coworkers are doing it, so by taking the time to understand and review LLM output you’re painting a target on your back? Anyway I’m not getting burned- the company is.


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: