NHacker Next
  • new
  • past
  • show
  • ask
  • show
  • jobs
  • submit
Smaller, faster, safer: running Kimi and GLM at scale (blog.cloudflare.com)
scrlk 1 hours ago [-]
Nice to see a provider being transparent about KV cache quantisation. I've been suspecting that some providers do this silently whilst heavily promoting their unquantised weights, even though KV quantisation can degrade quality more than weight quantisation.

However, I wish their testing were more detailed. Firstly, some model families are more sensitive to KV quantisation than others (only Kimi K2.6 was tested). Secondly, the evaluation suite they use to claim that FP8 KV quantisation is indistinguishable is noticeably lacking coding benchmarks; in long-running tasks, minor tool call errors compound over time.

anonova 1 minutes ago [-]
vLLM's study also concluded that "FP8 can deliver meaningful latency and capacity gains with small or negligible accuracy loss". Their benchmarks include LiveCodeBench 6.

https://vllm-project.github.io/2026/04/22/fp8-kvcache.html

syntaxing 41 minutes ago [-]
> View pricing in the Cloudflare dashboard ↗

Why… I wanted to see if it’s worth it to use cloudflare’s endpoint but I can’t even see the pricing

brokenodo 2 hours ago [-]
I was interested in reading this until my slop detector went off at the paragraph starting with “It's worth being precise about where the benefit comes from, because it isn't raw speed.”

I love AI, but I really hate reading it.

hankbond 1 hours ago [-]
I have had to stop commenting this because it would end up on 50% of the posts here. I really wish we could flag prose as ai-generated on here and just filter it out.
dgellow 19 minutes ago [-]
Don’t stop commenting about it, if there is something we (the readers) can do is ensure it is seen as uncool to post slop content
gr_norm 1 hours ago [-]
LinkedIn (of all places!) announced a button for flagging this recently: https://www.linkedin.com/posts/hsrinivasan1_ai-slop-is-a-top...

How well it would work on this site, I'm not sure.

speedgoose 51 minutes ago [-]
If it works, it’s going to be the best feature introduced by a social network in a long time. Incredible that it comes from LinkedIn.
Oras 5 minutes ago [-]
If there is an action on AI slop on LI, it will end up with almost no posts at all
hankbond 55 minutes ago [-]
Next up, LinkedIn starts using this feedback to train a classifier. They then announce an officially approved "not slop" classification only for LinkedIn Gold member posts. The classified posts have a wider reach due to everyone filtering out AI slop. Non-members automatically get bucketed in with the slop bc they don't pay to have the verified classifier run on them.
physix 52 minutes ago [-]
Better would have been to offer a button to flag something that does NOT seem like AI slop on LinkedIn.
trollbridge 53 minutes ago [-]
Sign up for Pangram and install the browser extension; covers X, Reddit, and Substack, and more to come.
arjie 56 minutes ago [-]
Cloudflare blogs are not meant to be human-read, AFAIK. They're raw material meant to be fed into an agent to be filtered down. I rarely read the contents because they are usually word-expanded to a greater degree than an article from The Atlantic.
dgellow 17 minutes ago [-]
That’s disappointing, in the past cloudflare had some of the best engineering blog articles
Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact
Rendered at 20:41:58 GMT+0000 (Coordinated Universal Time) with Vercel.