Start with how much the task has to read. Changing one behaviour across a dozen files wants a large window that holds up at length; editing a single function is faster and cheaper on a smaller model.
For long runs, consistency beats peak. A model that is brilliant sometimes and off the rails often just moves your time into cleanup. Run the same prompt three times and look at the variance instead of the leaderboard.
Budget per task, not per hour. Long runs re-read files, run tests and retry, so token use is not linear. Set a timeout and a step cap so it fails early.
Prices, context sizes and the model list change; take them from your account page on the day, not from this page. What stays fixed is the process: trial on a small task, check variance, cap the run, hand the diff to a human.