I have to spend hours understanding how other humans have asked the computer to do something, and they are also bad at asking it to do so, so I could have gotten a better result if they just came to me with their desire...
It's the same price as regular processing. You get guarantees microsoft give you, which are ones OpenAI won't (or require dedicated spend,) and you can use azure identities for access.
We use it for access to gitlab, ado, github, datadog, rancher. I think the rancher one is the worst. We develop custom MCP servers for our internal stuff for agents to use, and it all gets accessed thru agentgateway.
I don’t really like having to use MCP but we don’t have a good solution for authorizing individual calls outbound from a sandbox without choosing to just not care about the sandbox.
I would expect, although have no evidence, that any obviously high entropy crap like base64 and so on probably would get removed whether it's a secret or not.
I am running GLM 5.3 across 2x DGX Sparks and was doing comparisons and it absolutely can beat Gemini. Yesterday it corrected a poor Fable 5 response even
Yes they are quite good, but are not able to run on a 16GB RX 9070.
Quantized Qwen 3.8 Flash Next could maybe run eventually on that card with a highly optimized inference engine that dynamically caches the hottest layer experts. Even then you run into some hard limits.
reply