Hmm, perhaps I should switch to an MLX version… problem is, it took quite a bit of work to get llama-server (with llama.cpp) to serve my model(s) and allow requests in non-thinking (default) and thinking modes.
Thanks for sharing!
Did you observe a speed difference between ollamas mlx version and the mlx-community/Qwen3.8-27B-4bit from HF ran with mlx_vlm.generate (with MTP)? Or is it the same?
I'm not sure if I can get rid of the drafter model, if I understand correctly, the Qwen model already includes a built in draft headers, but just having --draft-kind mtp results in about 17 t/s.
The problem is it's not the engineers that overthink it. The requirement for soc2 usually comes with the first "serious" customer. It is usually a big blocker on some fat contract and now the business makes it your problem for the next 6 months.
So what do you do? You engage and some 3rd party 1800-need-soc2 clowns which will hold your hand and implement all the cookie cutter solutions they know will make auditor happy (oh and btw, they know the auditor personally).
No, that's not a brag but a suggestion to take things into context and not apply soc2 as a cookie cutter solution where a 20k employee enterprise and a 20 people startup must share commonalities.
In a 20 people startup it's very likely that most engineers have access to production anyway and can inject malicious stuff directly, so PRs stop no one really.
This is a pretty harsh critique, while I see the agent related markdown the application seems measured and well applied. Hardly seems like slop; it's might be a one man show+ai agent but the direction and project composition feel like it's something that's logical and manageable.
I personally like the idea even as a p.o.c (inline with bellards work). I'd like to see how nostr & cloudflare workers providing out of band ops on what is effectively a container per tab.
Show dead and you see 23 posts about this exact same thing which gives me pause and makes me wonder, did this guy really post this 23 times?
Why? Or is there some program setup on a loop?
Either way,a legit Linux shell in a browser would be kinda powerful and even if this one isn't legit, Im intrested in adjacent projects...for instance this may be the only way to get a shell in a usable shell on iphone which would be neat and useful