Public shared posts

riomari shared this post · 1d ago
Wong Hao Shan

LLMs are getting smarter, but they still generate one token at a time. 🐢

For agentic workloads involving long reasoning, tool calls, and code generation, that sequential process results in higher latency, higher compute costs, and lower GPU efficiency.

Speculative decoding helps by letting a smaller drafter propose several tokens for the target model to verify together. DFlash went further by drafting an entire block in parallel.

The remaining challenge is coherence. Each prediction may look right on its own, but the complete sequence may not fit together.

96
Anvesh Gummala Diffusion text models
Diffusion Gemma 4
2d ago
Muhammad Usman Very interesting must read! 4d ago
riomari shared this post · Jul 8
riomari shared this post · Jul 8
Athharv N.

If only I could keyframe my actual tennis skills this perfectly. 🫠

Just survived my first 6 months on the court. Naturally, I had to turn the obsession into a motion loop.

Turn the sound up for this one. 🔊

Design, Animation, and Sound: Athharv N.
Tools: Illustrator, After Effects, Cavalry App

#MotionDesign #2DAnimation #SoundDesign #AfterEffects #CavalryApp #TennisLife

241
Aniket Sawate Craaazzzyyyyy🔥🔥 Jul 8
Robert Johnson This is amazing! Great work. Jul 8
riomari shared this post · Jul 8