The 996 working culture in this page is actually *996.ICU*. "Work by '996', sick in ICU", an ironic saying among Chinese developers, which means that by following the "996" work schedule (9 am to 9 pm, six days a week), you are risking yourself getting into the ICU (Intensive Care Unit). [1]
Take a step further, do you want to overwork to death if you have other choices? One of the reasons behind 277k stars in 996.icu repo is that 996 working culture literally takes over every single corporations in China. With the downturn of China economy and the selective execution of Chinese labor law, in recent years, 996 becomes even harsher to labor force in China because there are more than overwork now, e.g. low salary and low human rights. When you say 'make the average consumer better off', you conveniently ignore the tears, sweets, and blood behind the facade of neijuan.
> I was shown a prompt and told the internal research model had simply been given the problem statement. Levent had been told by Sebastien “very little human input” had been used. This turned out not to be true. Over the course of the call, as members of their team sent Sebastien corrections and details over their internal chat, it emerged that an entire team had been working on the problem, that this was one of a number of things that was tried, that work had started on the unforced problem, that the team first set the model on easier problems, including Euler, that even the prompt that had been shown to me had been written by prompting Codex, and that an insane amount of compute had been used.
> I asked when the first prompt had been sent by them. This question was not answered directly by OpenAI for some time. Eventually it was agreed that it had been sent in the past few days, after information about our work had reached OpenAI.
> I asked whether the model had been trained on, or had access to, our sessions in Codex, into which we had been putting all our drafts for the whole of this project. I was told the model did not look up user data. I asked again, about training, and I did not get an answer.
> Two proposals were offered to me. The first was that we post our Euler result, and that OpenAI post its Navier-Stokes result the next day. The second was that, after posting Euler, I alone write a paper presenting the Navier-Stokes result, acknowledging that an internal OpenAI model had resolved it. Sebastien twice asserted that he wanted Levent removed from authorship, and said it would all be simple if only it were not the case that, and it was so annoying that, Levent works at Anthropic. It was also said that if OpenAI posted after us, they would say that we deserved the Clay Prize, and that we were the “closest humans to the problem”. I declined both offers.
> I said that if OpenAI released its result in the way proposed I would go public with what happened. The reply was, “Why would you ruin your career?” I replied that I am an academic, and asked why he thought going public would ruin my career. The reply was, “If you don’t want me to be nice, then I don’t have to be nice.”
Wow, that's some VERY friendly communication. Besides, will the career of the person be ruined because of “Why would you ruin your career?” came out of his or her own mouth?
For people who want to quantize kv caches for longer context, basically, if the model checkpoint you use didn't trained to use quantized kv caches (and all most all of model checkpoints you can access didn't), you shouldn't quantize kv caches during inference. A major benefit to do quantization-aware training/distillation on official checkpoints is to let the model to be familiar with quantized activations/weights/kv caches.
In contrast to memory-bounded autoregressive decoding, diffusion decoding can be computation-bounded. Devices with a decent amount of compute but very small memory, e.g. nvidia consumer-grade GPUs, could benefit from diffusion decoding massively.
Performance-wise, I don't think diffusion decoding should be much worse than autoregressive decoding, see https://arxiv.org/html/2604.11035 .
As a local LLM user, I really want to see more small diffusion LLM/VLMs with decent performance coming.
Compared to closed weight (especially unreleased and access-limited) and open weight/source large sparse MoE LLMs/VLMs, open weight/source small dense models benefits public the most because they just reaches more people.
Compared to Qwen 3.6, 3.8's thinking style changed drastically. With xhigh budget, it thinks a lot MORE, and longer thinking session directly translates to better performance. This tradeoff between performance and computation, memory, etc. is meaningful to me.
However, because Qwen 3.6 and 3.8 share the same architecture, with 32GB vram, llama.cpp, IQ4_XS model, MTP and FP16 mmproj, I can only get 200k context, which is not good compared to 640k context of muse glimmer. Hopefully this problem will be solved in Qwen 4.0 release.
Now open weight LLMs/VLMs/LMMs are becoming even larger to the extent that consumer-grade hardware are no longer able to run these models. In contrast, quantization and pruning make the model better at the size-performance pareto and provide people with strictly more possibilities.
Other details for the official (Yang Youlin) in this news.
---
Whistleblower Yang Hai already reported Yang Youlin for his economic misconduct in July 2008. The whistleblower was detained because of the report at Nov 21st, 2008.
---
People's comment on this matter around March, 2009:
哇塞!终于有人敢动杨友林啦,杨海好样的。杨友林此人早该除无奈碍于他的势力。除掉杨友林大快人心。Wow! Finally, someone dares to take action against Yang Youlin. Good job, Yang Hai. That guy should have been dealt with long ago, but we were stuck with his influence. Getting rid of him is incredibly satisfying.
不杀此贪官,难平民愤。If you don't execute this corrupt official, you won't appease the public's anger.
江宁有一个传说,谁也动不了杨永林!There's a legend in Jiangning: nobody can touch Yang Youlin!
他的保护伞是谁?Who is protecting him?
希望引起中央的重视!I hope this gets the attention of the central government!
现在社会怎么啦?好多天了根本没人关注这件事?是上层没有看到?还是视而不见?还是怕牵连自己?What's going on with society lately? It's been days and no one is paying attention to this! Did the higher-ups miss it? Are they turning a blind eye? Or are they just afraid of getting dragged into it?
---
Since the report, there were several pieces of news about "the investigation about Yang Youlin is ongoing", but no real progress until 2023.
Basically, Microsoft logs things from windows users, including and not limiting to, the machine GDID, IPs that come with the GDID, and when and what exact URLs accessed. So for windows users, privacy, an important part of information security, is totally destroyed by these logs enforced on windows. Another important part of information security is bug fixing, and microsoft did make at least one security researcher angry [1].
And a simple solution to that problem is moving to linux. You save yourself a lot of time and energy for leaving the adversarial information security condition imposed by windows and microsoft.
P.S. You may consider debloating windows for a more information security friendly environment. However, that is nearly impossible, as long as you realize that windows is an OS composed of thousands of closed source softwares, and doing security audits on all of these will be costly.
I've been using searxng for several years now. I don't run my own instances because the inhumane network censorship imposed by GFW, and proxy detection enforced by search engines. Instead, I rely on public instances on the list [1] and libredirect [2]. Note that service from a single instance is not guaranteed, but you can always switch to other available instances with little cost within a minute.
I won't say searxng can help you degoogle because metasearch engine calls other search engines, e.g., google, to collect results. However, if you try searxng, you can at least get rid of things like ai reviews in no time.
In the end, thank you people after searxng project and public instances.
The thing about the public instances, is now you often have to go through a lot of them to verify they work properly. SearXNG needs better quality control.
Often have to go through the preferences to deselect search engines that don't work (often because of the instance being blocked) or select those that do work, because of reliability problems. Which engines are working, can be different for each public instance, so that even saving a preference hash doesn't always work.
Would be great if SearXNG did automatic adjustment of presented search engines (or offered the option) based on reliability.
Yeah, searxng instance reliability problem goes deeper than simple uptime. The google response time function on public instance page [1], a good measure of the search engine availability of one instance, is broken for some time. Without real testing, people really cannot get the full view of the service status of a single instance.
The 996 working culture in this page is actually *996.ICU*. "Work by '996', sick in ICU", an ironic saying among Chinese developers, which means that by following the "996" work schedule (9 am to 9 pm, six days a week), you are risking yourself getting into the ICU (Intensive Care Unit). [1]
Take a step further, do you want to overwork to death if you have other choices? One of the reasons behind 277k stars in 996.icu repo is that 996 working culture literally takes over every single corporations in China. With the downturn of China economy and the selective execution of Chinese labor law, in recent years, 996 becomes even harsher to labor force in China because there are more than overwork now, e.g. low salary and low human rights. When you say 'make the average consumer better off', you conveniently ignore the tears, sweets, and blood behind the facade of neijuan.
[1] https://github.com/996icu/996.ICU
reply