If you buy tickets online from AMC, they now show a disclaimer that tells you the movie starts around 25 minutes after the listed show time, though usually it starts 30+ minutes after.
It's not just web chat, V4 Flash 7/31 suffers from a lot of pathological behavior in coding harnesses as well, e.g. infinite loops, hallucinations, premature termination, and invalid tool calls.
All of these flash models have this. You have to build your harness so that it deals with it. Infinite loops are solved by having an error message that says what to do differently on failure, invalid tool calls are solved by making the tool schema less strict and detect things in the runtime etc.
Hallucinations you can't fix. Gemini is a bit worse there than DeepSeek, but there's not much research on how to fix that. The only one is the CaMeL paper by Google, where you tag every prompt and result and then for every assistant response or tool call you first check where it got that data and error if you notice fabrication. This one is really annoying to implement.
With larger models the fabrication starts when the context grows or if you have too many tools, for flash models it's much earlier. We use the flash models for repetitive agentic tasks, where the prompt defines clearly what to do and how. The whole run is about 4-5 steps typically, and context size stays in the comfort zone.
Can you share what tools and processes you're using to do this?
I've been using Pi to build custom extensions and wrapping workflows in shell processes to make it more deterministic and enforce certain validations, all guided by Fable. This isn't production work, though, just playing llm factorio at home.
What you want is a bunch of sessions to replay. Something anonymized if it's not yours, and something that's not depending on state.
You replay all your sessions against your harness, and then store all logs all output, everything to a safe place.
Finally use a blind judge to check everything, and score the output.
Then fix your harness, iterate again until better until you are in a point where it's just the model's weakness. If you get to that, use a bigger model.
Same. My side projects are coded almost exclusively with the Deepseek V4 Flash 07/31 in omp, and it recovers beautifully in every case. I'm using OpenCode Zen.
OpenCode but I variously use DeepSeek API/OpenRouter/Vercel AI gateway. I'm sure it's the combo of model + inference provider that is the issue and not the model alone. DeepSeek API also has far better inference speed and reliability than the cheapest providers. That said I never seem to have these issues when using GLM 5.3 flash served by OpenRouter/Vercel.
disagree; been using flash as my exclusive model (other contributors have used other models) to build a complicated software project, a web engine. See https://github.com/gterzian/formal-web, which as you can see comes with very specific guidance explaining how to implement features.
I'm using headless Pi with my own UI and sandbox client, https://github.com/gterzian/uni03C0, as well as a bunch of Pi extensions for things like accessing Web standards and browser use via CDP for testing.
Switching to 4.1 today...
Edit: it seems they pushed the date at which they route the Pro calls to new Flash, so today I ended up paying regular Pro rates thinking I was using the new Flash; an example of how their offering is not quite as predictable as I would like it to be (the other is cache performance being unpredictable).
I have used the flash model for over 3b tokens and ofc. I saw some hallucinations and premature termination (I also get this on Astra - way more often than with deepseek v4 flash), but I never had a infinite loop (using the copilot as harness).
I used a lot V4 flash to implement plans built by other models, and it was honestly top notch. The thing was a workhorse, and I got none of the isuses you describe.
I was mostly using DeepSeek on Pi, connecting to their API directly (not some third party provider).
I honestly have more issues steering Sonnet properly.
Well, for one, most math proofs don't have any practical applications, so a proof that no one reads is basically a digital paperweight. You might as well suggest AI write novels for other AI to read.
The hope is that some of them end up being useful; otherwise, nobody would be funding math departments. Mathematics typically anticipates and enables new physics and chemistry.
If people are just doing math to kill time, I don't get why anyone would bother with AI. Do people really enjoy picking through a million lines of generated Lean code, if it's not for any practical use?
If you're interested in the topic enough to comment on it, you'll probably find it worthwhile reading a mathematician's perspective. Here's the prolific Terry Tao:
https://mathstodon.xyz/@tao/117219548485446992
They don't actually say anything about why anyone should fund this, though. I don't get why a society should worry about progress in mathematics if there's no practical benefit expected.
Maybe there's two kinds of math that we need? Useful math and navel gazing, and we can hand the first to the machines, and let hobbyists do the second in their free to entertain themselves?
It's hard to know what math is 'useful' a priori. That's always been the argument for supporting basic research. This is not why I am a mathematician however. I think there's intrinsic value into understanding something of depth and meaning, but the societal setup we have now that mostly agrees this is valuable is probably a very contingent phenomenon that is unlikely to last much longer.
Yes, so if there's useful math, you throw the LLM at it and use the results, no humans needed.
Humans can try to extract some ideas from the million line lean proofs, if they want to, I guess. But I can't imagine anyone really funding the human part of it.
A lot of movies gave the impression that making phone calls and balancing checkbooks would be the easiest tasks for consumer AI to solve, while math and science might require exceedingly advanced AI. Turns out to be the opposite: computers are great at math and bad at conversation.
Who complains about sync thought? What tilts me is Pocket, AI features, integrated ads, and a plethora of other random acquisitions and side projects that constitute outsized expenses yet do zilch to drive adoption.
I also used to be a Thunderbird lover, but I was forced to ditch it due to crippling freezes, crashes, and slowdown when you have a lot of emails. Mozilla just feels like an awful steward of their projects at times.
EDIT: I also subscribed to Mozilla VPN for almost two years and was forced to ditch that too, due to... solvable technical issues, again. Mullvad was just a greatly superior product for the same price. I still use Firefox for all my browsing on desktop and Android at least.
Totally agree they lose focus and spend a lot of time on things that are questionable. But again, they do release things that are super useful. There are probably people out there who like pocket. As for AI, as much as I was raising my eyebrow, I watched a whole threat of people the other day raving about Kagi’s AI implementation. Maybe Mozilla will make something useful shrug
1. Pocket was an external company, which Mozilla bought (and claimed they would open source), which already had a web extension. There was no need for the pre-acquisition integration, or to buy it later on.
2. After buying it, Mozilla did nothing with it, including failing to open source it, and seemingly failed to get enough revenue from that they shut it down last year.
This saga provided no benefit to Firefox users who did not use pocket, and at best for pocket users left them with a dying solution that they would need to migrate from (given the failure to open source it). It's not clear how much Mozilla paid for Pocket, but those funds could have been spent on something more valuable.
As interesting as all this is, I still feel like the threat model for "unconstrained black hat AI agent cluster" is probably weaker than that of "highly infections network virus" because it is much harder for an AI agent to hide or replicate itself at this time. Maybe the day comes that it takes less than an 8x GPU node to run a state-of-the-art LLM and the risk of SkyNet increases. For now the potential for intentional cyber attacks feels like a much bigger threat than accidental hacks. (That said I have little cybersecurity background.)
There is a section in the article answering your question if you read it.
> Suddenly, the crawlers were coming from millions of random residential or mobile IPs, all pretending to be random modern browsers. An IP like that would make 4-5 requests and then never show up in the logs again. There was no point in banning them, because by the time you figured out that they were bots, they were already done with you.
I’d love a service like spamcop.net where I could submit my access_log and they lookup the abuse addresses and file abuse reports in my name. Maybe if people’s Internet access gets suspended they’ll think about installing random apps that work as a proxy in the background.
That is a ridiculous way to try and deal with the problem of residential proxies.
You are, in reality, only hurting the actual owners, the subscribers of those ISPs who are behind those addresses. We call that "collateral damage".
If any of those actual residential users try to use a website, their ability to freely access the Internet may be harmed by a bad reputation that they do not deserve. They may be totally unaware and non-consenting to residential proxy use.
You are not, in fact, hurting the residential proxy-ers at all. Not one bit. They will move on to another IP and another compromised LAN, and they will continue to move on and on and on. They will not be harmed or impeded; they will simply keep turning up fresh, new, high-reputation IPv4 and IPv6 sources. This is a sheer numbers game, where the numbers are always in favor of the attackers.
Also if network admins keep blocking/filtering abusive residential proxies, they will balloon their firewall rules and cause actual performance issues at the network level. You will turn into your own DDOS without any actual benefit. You're on the losing side of the numbers game, and in the immortal words of W.O.P.R., "The Only Winning Move Is... Not to Play."
Similar to how people running an open SMTP are complicit in promoting spam, I see people running a wild public proxy as complicit in this malicious scraping activity.
And similar to how most people running mail daemons are using blackhole lists nowadays and are keen to not end up on there, maybe ISPs and web hosters can use the AbuseIPDB to sort out their customers.
Just doing nothing doesn't appear to stop the scans hammering my poor Raspberry Pi serving my few Git repositories.
Hey, from the beginning of SMTP, running an open relay was an administrative mistake. The MTA administrators were supposed to know what they were doing, because resources were allocated to them. They had privileges granted for the system and the network. It was right if they were blacklisted for misuse of those resources.
Now in 2026, running a "public proxy" doesn't take an administrator. You don't even need to be aware. Most victims are unknowing victims. They simply subscribe to an ISP and they have their own devices. They are being exploited for that innocence and ignorance. Most victims have no visibility to even detect that they're being used as a proxy. Most victims couldn't stop it, even if they wanted to.
I challenge anyone with a home router to list the processes running on that router, and list all current open connections on that router, and list all open, listening sockets on that router. I bet you can't do it. There are no consumer router OS that lend themselves to being secured, or even diagnosed. Malware can easily be planted on any of them and run, completely invisibly.
A residential proxy server could run on routers, could run on a switch, could run on your "Smart TV" or a smartphone, or a notebook computer. It could be anywhere in any form. Perhaps you consented to it, perhaps you didn't notice.
In no way is this the same as an SMTP open relay situation. If you wanna play "whack-a-mole" with a "blackhole list" you're simply going to overwhelm those lists with false positives and collateral damage. The residential proxies have long since moved on. You won't even find the culprits using those addresses you just blocked. You're just clogging up your own machines. It's a total self-own.
> You are, in reality, only hurting the actual owners
How many times do I have to hurt them before they decide to buy a different smart TV?
Seriously, that's like saying "if you try to stop your neighborhood rodent problem by getting citations sent to people with cat food on their porch, you're just hurting the innocent outdoor cat owners". They're participating, whether they know it or not. We can and should PSA and shame and regulate away residential proxies on the supplier side, but we can and should also simultaneously discourage them on the end-user side as well.
Comparing it to Alzheimer's is an allegory. The cognitive loss is already well documented. I might also compare it to Chronic traumatic encephalopathy. But I think more people understand what Alzheimer's is.
I argued this idea to a couple of my classmates when I was a physics undergrad, and they agreed. However, I later changed opinions because of what this does to the derivatives/integrals of your trig functions.
For general periodic functions, [0, 1) is a good domain. But circles and spheres are geometric objects, and radians/steradians are geometrically significant units that are well suited for general purposes.
I do remember that Doom uses an interesting alternative representation where an angle is a u16 multiple of `(2 * pi) / 65536`. Fixed point is sometimes a good choice in games and simulations due to having uniform precision.
I've seen many accusations and exposes of clothing brands secretly sourcing sweat shop labor, but this is the first time I've seen anyone allege that sellers are secretly hiding the fact their clothing is made in high-tech ethical factories.
Me too. My understanding is that clothing manufacturing is basically impossible to automate with our current level of technology, due to the difficulty of manipulating fabric.
I'm sure they can't automate shirts because the super cheap ones somehow always are Made in Bangladesh. If they could use less human labor they would make them somewhere else.
I believe socks can be fully automated and are mostly made in china.
There is a startup called Unspun that wants to make many types of clothing production fully automated by weaving the complete end product in place. Similar to how socks are woven today.
The current way to make a "finished" clothing product is called "whole garment knitting". Only two companies make machines that can do this; designing the garments is a pain with the current software stack. Most clothing is "cut and sew" -- you make a simple sheet of knitted material and cut and sew a garment out of that sheet. It is wasteful and heavily reliant on cheap labor.
reply