Hacker Newsnew | past | comments | ask | show | jobs | submit | bitexploder's commentslogin

For what it’s worth, I am very happy with Jellyfin and the *arr suite. It took a bit of agent prodding to get them all playing nicely together and bypassing Cloudflare CAPTCHAs. However, it's pretty sweet when you get it all working.

> bypassing Cloudflare CAPTCHAs

Can you say more about this? I've never encountered this problem in nearly a decade of running the same stack.


With the *arr stack, some of the trackers it uses to search for a torrent will present a captcha when Prowlerr queries it for a search. If you integrate FlareSolverr and set up Prowlerr to submit queries through it for the trackers that have CAPTCHAs, it won't get blocked by them.

It does seem like for new code that might help. There's some really good logic and wisdom in it, but it has to be applied very contextually to the exact problem you are trying to solve. If an agent is navigating a complex codebase, this could definitely send them off on a refactoring rabbit hole. However, if you have them writing some new code, it could prevent their tendency to yak shave and write new things. So I can see some situational uses for this, but it could get out of hand as well.

Sol is my current favorite model to interact with. So much less BS than Opus 5. Fable 5.1 is okay as is Fable 5 but it has Opus like tendencies. Sol is very good at following instructions and remembering them for a session.

I still hold the line on interviewing. Maybe some don’t. I am sure it is true. My team us too small with too much responsibility to tolerate mediocrity to any real extent

In modern America the answer to that question is often resoundingly yes. Not just hypothetical.

Opus 5 is a genuinely infuriating model. I hate it’s behavior.

We have a mandate to only use 4-8. Sonnet 5 is pretty good and fable 5-1 has been pretty good so far fwiw.

Fable is okay, just slower, eats tokens and not any better at coding tasks. Maybe a little better, but not better enough. It's a lot faster to have a cheap and fast flash agent / sonnet do the implementation work with Fable tagging cleanup and divergence from spec and goals.

Flash 3.8 is genuinely my favorite all around model right now. And yeah Opus 4.6 was the last Opus model I liked. 4.8 is tolerable. Opus 5 is a terrorist. It just can't follow an instruction to save its life and regresses rapidly. Sol at least stays on track so I have to smack it's hand way less often. I am biased, but Flash 3.8 and 3.7 are the first Gemini models I just recommend to others.


5-1 is considerable cheaper I’ve found and feels a lot more like 4-6.

But opus 5 described as a terrorist is being generous.


Amusingly, as an autonomous coding agent, I kind of like Opus 5. But I have to bound it on tasks or it just goes off the rails. But I'm bounded tasks, it is genuinely solid. It's kind of like the new Sonnet 5. Right now my favorite model to interact with on the frontier side is Sol 5.6 so I have been using that as my coordinator. Flash 3.8 is my other favorite just because it is so fast and I use it a lot at work and know its quirks.

I feel like a lot happened this week and people are glazing how ridiculously strong Flash 3.8 is right now compared to Fable/Opus/Sol/Astra.

Flash 3.8 is rad. Easily my daily driver now. Only downside is it's Gemini so sometimes it just keeps going until it wants to be done.

I have a few attention and finish mechanisms in my prompts. I have been using it for a week and a half and with some prompt taming it is great. (I have early access to the models cause I work at the place that makes the model). None of my attempts to ever tame Opus 5 have worked.

I wish people could see how some of this reads. You are an “amateur” using a model 6-8 weeks behind? Really? Sigh.

Does Claude not allow third party harnesses?

The thing I didn’t realize for a while is 27B is rather smart. As many (or more) activated parameters as the flash models of the universe that we know about. It reasons very well. It just doesn’t have a lot of knowledge.

They seem to have good enough general intelligence that missing knowledge is not that big thing. If you are able to have a proper [free search engine], they can do almost anything. Having own local search index about relevant stuff can help a lof if you don’t want to pay for search API.

But running that fast… with a local RAG? Yeah, it is a very interesting model. Maybe you don’t need a lot of parameters, just a really big local database :)

You can run it with 2x r9700 with 150-200 tokens per second. It is intelligent enough if you just point the docs / whatever for it.

I believe. I run it on my mac M5 pro at like 30t/s with some RAGs and let it work on stuff overnight and it's great. It isn't the same as the big models where things can be more unbounded, but if local models keep progressing there is a universe where a 200-300B model is all most of us will need to stay out of the big tech moats.

Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: