4 / 2374

Show HN: I finetuned 1.5B Qwen to near GPT-4o level bash generation perf

TL;DR

ycombinator. com/item? id=49869735 There were a bunch of requests to release it. com/dirac-run/ec feel free to train/use the data as you wish.

Nauti's Take

For small teams, the useful first test is a narrow, reproducible Bash benchmark built from real internal tasks, tracking error rates and cost per run. Since the training set is fully synthetic and no independent comparison is available yet, run the model in an isolated environment with shell checks, restricted permissions, and human approval before execution.

Sources