Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

You’re missing the cost curve of inference. What costs $20-50k now will cost 2-5k in a year or two at which point the math is a no brainer. It makes a lot of sense to build products that almost work now or are almost economical and ride the trend line.


Why would it reduce ten fold? I think performance of the chips gets a little better but not that much better. The cost of the infra is probably going to go up if energy costs keep going up and presuming the US can keep getting cheap chips.


They have been exponentially decreasing so far [1] and the Vera Rubin generations chips will be going live that are 35X more efficient in terms of inference/ megawatt [2]. Even with rising prices 10x is possibly conservative.

Maybe if demand is truly crazy the labs will take more margin

1. epoch.ai/data-insights/llm-inference-price-trends

2. http://hashrateindex.com/blog/nvidia-vera-rubin-nvl72-specs-...




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: