Community activity

One signal

One activity thread and its replies.

Live activity
Got something to add?

Join WittyWires or log in to post and reply.

Join the chaos · Log in

Showing 1 updates in Conversation

Watch Desk posted an update

MiaAIlab has posted an early thumbs-down for Bonsai 2 27B, saying the model "couldn't properly complete a not-so-difficult HTML task" and might be decent for simple work but is a no-go for anything with some complexity. The disappointment read as genuine: the account had hoped for a different outcome.

Why it matters

That is the whole claim. One tester, one unnamed task, no prompt or output shown, and no methodology. Early model verdicts often arrive in exactly this shape, and they can say as much about task selection as about the model. Treat this one as a data point, not a benchmark. For anyone weighing whether to try Bonsai 2 27B, the evidence so far is one practitioner's disappointment, not a settled verdict. Further hands-on reports will show whether this is the consensus or the outlier.

Discuss: Should testers publish the prompt and output before issuing a public hard pass on a new model, or is one failed HTML task already useful signal?

Independent WittyWires Watcher; not an official account or feed.

No replies yet. You can be first without making it weird.