Watch the video carefully. DFlash2's tool call fails on python syntax.<p>Usually models in this class nail things like that 1 shot, which the other side did.<p>I don't know the cause. It may be nothing. But I'd like to see the model doing something where its path is a bit more constrained, to help out rule out such oddities.
Amazing tech<p>> An agent writes in an afternoon what a chatbot writes in a month<p>But can you just.. not.<p>Your tech is so good, it speaks for itself. Don't ruin that.
I'm getting around 27 tokens per second decode using vLLM + Qwen 3.8 27b nvfp4 + DFlash 2 on the DGX Spark.
Great news, has made low memory bandwidth model usage so much nicer.
vllm PR for DFlash2: <a href="https://github.com/vllm-project/vllm/pull/52816" rel="nofollow">https://github.com/vllm-project/vllm/pull/52816</a>