7 Comments
User's avatar
Andrew Sniderman 🕷️'s avatar

On a slightly more serious note. I sat in on a talk on SDLC (software development lifecycle dontchaknow) and it reminded me of the progression from waterfall->agile->devops where now we ship continuously and you (mr consumer) dont gaf.

When does that happen in this space? I think once consumer settles (agentic agents everywhere?) model pickers go away and it’s just plumbing. MSFT is pushing for a general biz layer where model picker is optimized for workload / cost.

We’ll for sure still have special purpose models for fields medicine, law, etc. If/when this comes to pass, I wonder what happens to the frontier labs? When we no longer obsess over the latest model do they disappear as plumbing or will they grow into something more?

Looking at OpenAIs devday announcements, they are going for a fat middle layer of the stack ala MSFT back in the 2000s

Daniel Nest's avatar

Once AGI hits in 2027, there'll be no “frontier labs” - it's all AI agents spawning AI agents spawning AI agents spawning paperclips.

Andrew Sniderman 🕷️'s avatar

Sagrada Família’s lawyers would like a word

Daniel Nest's avatar

Joke's on them, they're all dead by now. Thanks for the long-term planning, Gaudi!

Andrew Smith's avatar

I've got several ongoing projects where I'll hit a wall with the current model, then the new one comes out. That's another pretty easy way to benchmark, but it only works if are crazy like me and have way too many unfinished things.

Daniel Nest's avatar

Would be curious to hear if this works for you!

The idea is that it should also be good for recurring, well-scoped tasks that can be done better with better models, so see if that's something worth trying.

Andrew Smith's avatar

That's the thing: these projects are big, so there's a lot of recurring work. It's like an ideal test.