Hacker Newsnew | past | comments | ask | show | jobs | submit | kroaton's commentslogin

Garbage.

I think it mostly shows that there is no moat and the only advantage the U.S companies have over the Chinese is more compute. Qwen Max, Kimi K3, GLM 5.3 are really close to Opus/Sol/Fable/Astra and they are open weights.

  I think it mostly shows that there is no moat
You can argue that TSMC has no moat since Intel and Samsung are also able to eventually make a node as good as TSMC - just a few years later and at smaller scale.

And no one would say that about TSMC.

So there is clearly a moat there somewhere.


No. In the semiconductor industry, the "catch-up" player isn't normally spending less in absolute R&D terms.

Comparing the R&D costs of creating GPT-4o vs. DeepSeek V3 (the latest gen for which we already have good accurate numbers) it looks like the latter cost 1/20th as much to create.

If Samsung could catch up with TSMC for 1/20th of the cost, people definitely would say that TSMC has no moat.


Why do you think Chinese models cost 1/20th to train?

That's the ratio the widely published numbers give [1]. One does not have to believe the numbers [2], but those who do believe them are then justified to conclude that there's no moat.

Which numbers you believe is of course going to affect whether you think there's a moat or not. That's largely orthogonal to your TSMC/Samsung analogy I responded to. If you think the "moatists" are wrong because they believe the wrong numbers, that's fine, but then there's no need for the analogy.

[1] https://galileo.ai/blog/llm-model-training-cost

[2] https://medium.com/@theiand/how-can-deepseek-a-5-6-million-l...


But fundamentally, why is their cost 1/20 and is it sustainable in the next 10 years of competition?

Now that is a good and interesting question! Hopefully a "no-moatist" will share their reasoning.

Because they're distilling frontier models and that's a lot faster and cheaper than training a frontier model from scratch?

So why can't OpenAI/Anthropic also distill the good parts of free Chinese models? It's even better and easier for OpenAI and Anthropic. No poison pills as well.

Ultimately, that's what I need to be convinced. No one has put forth a good argument yet.

Clever architecture --> Ok but OpenAI/Anthropic can use these as well and they also have very smart people with their secret clever architectures

Distilling --> Ok but distilling means you will never be smarter than the original. Furthermore, reasoning is now hidden by private labs and they have poison pill answers for distilling if they can detect it. They will be able to detect distilling better and better.

Cheaper electricity --> Ok this is cancelled out by their chips being much less efficient due to not having ASML EUV machine access.

So I don't see why fundamentally their training costs are cheaper over the long term.

I'm looking for a no-moatist to convince me.


Labor. Smart labor would be much cheaper I'd reckon in China than in the US.

How much advantage in costs? What % of labor is training cost?

Mercor, Tacit Labs, Handshake AI... I suspect companies like these play a big part in model improvements, generating high quality benchmark/task-focused data for training.

However, these do require educated, white collar, workers.


Considering that frontier scientists and engineers in the US are currently taking home seven (or even eight, in some cases) figure salaries - pretty high, I'd reckon.

Would like to see the math since the claim is made.

I am not a "no-moatist" per se but one can argue their might be a plateau to how good a inference llm can become. If this is the case the playing field shifts to context, tools and harness, which are much cheaper to build an compete on.

Please just say what you want to say.

Yeah, I'm not sure if "no moat" analogy stands for chip manufacturing. Even if foundries acquire lithographic nodes, the procedures (temperature, duration, etc) are for them to figure out and are usually kept secret. This secret could be the "moat" that differentiates each foundry's operational capabilities.

bringing the price down b.c. competition != no moat.

There's not 100 frontier labs, it's not like airline companies


About the same, 5-10, when you consider major (aka frontier) airlines.

Actually not a bad comparison. Both burn massive amounts of up front capital to protect an oligopoly in the hopes their commodity product eventually pays off.


From my experience with complex coding tasks (AI infra), I don't think these open weight models are close.

Not sure if I agree, I tried GLM5.3 and it was pretty decent. Ok, it's not Opus, but maybe it's Sonnet?

The "moat" is the "harness", the app.

For most people, the app IS the AI.

And even for its wonkiness, ChatGPT has had the best UX/UI of them all.

The way to win the AI wars in the eyes of the common folk is through the frontend, to be the Apple of AI, as it were.


this basically says you don't believe there is real AI.

Read the second line guy

There are people all over the world who have no computer skills but they use ChatGPT on their phones daily

They don't know/care shit about models and all that

For them, if the app sucks, the AI sucks.


they don't have moat in hardware either

Chinese counterpart like CXMT and Huawei is begin producing their own chip

You cant block an entire nation level effort with tariff


I think the moat that China has is energy costs. It's taking learnings from the Bitter Lesson. If you role up scale and compute to the next level, it's energy resources. China has it and sharing open weight models is an effective means of removing the tech moat. This idea has been floating around for a bit now (I'm not taking credit for it).

It's not energy costs. The US produces about 70% more electricity per capita. Chinese households do pay less than half what US households pay for electricity, but that's because the NDRC sets prices below costs for households. They make it up by charging industry more, and the industrial electricity prices in China are roughly 34% higher than in the US.

> The US produces about 70% more electricity per capita.

And consumers use 4x as much per capita. Industrial generation per capita China comes out ~2x

> industrial electricity prices in China are roughly 34% higher than in the US

For which industrial customer and where? Chinese compute hubs are on par to slightly cheaper on pure electricity costs.

Conversely the US makes it more expensive with interconnect and upgrade fees as well as hefty take or pay contracts.

A 1GW datacenter in VA for example would add 5-10c kWh and a 12 year take or pay deal


per capita seems the wrong metric given the difference in population sizes and America's wealth. They have roughly 4x the people and have added 10x new power capacity in the last 10 years, not to mention lapping us in renewable and long distance transmission lines added.

They also benefit from the commodification of software/knowledge work since they own manufacturing

If there was no moat, nvidia and meta would have SoTA models too.

Nvidia does have one of the best completely open models. Open weights are nice but Nemotron is open training data too.

It is not in nvidia’s interest to be too good at model creation

But it is in their interest that their customers can use their models as a base for post-training and LoRAs.

They don’t necessarily need their own models for that

They have models for that. That's what the Nemotron series is. Not just open weights but open training data too and full tutorials on how to use them to fine tune or train your own models.

They exist to keep people using and advancing the tools on their hardware.


Why not? Commoditize your complement, and all that.

And if they get too good, they risk harming or otherwise killing their golden geese (their customers), who they are heavily invested in.

How? Imagine an open-weight model comes out that is somehow better than proprietary solutions. Now the marginal cost for the consumer is just the cost of renting the inference hardware, without having to pay the overhead of the owner of a proprietary model. And because it is cheaper, more customers want to use it, and Nvidia will sell the providers the inference hardware that they need.

1. No open ai and anthropic means no buying gpus to train. Now nvidia spends money on hardware training their own models. Opportunity cost plus expense.

2. Any open models created from this will not necessarily need their silicon, see apple mlx.


1. I don’t think that’s a very strong argument. OpenAI and Anthropic don’t buy the vast majority of GPUs they use they rent capacity.

Nvidia could just the same rent those GPUs out for inference and actually have way better margins than they do right now. Antitrust and putting all your eggs in one basket are why they don’t, similar to TSMC.

2. Neither do AI labs. See Anthropic buying TPUs, deploying with AMD. OpenAI on Maia, Cerebras, their own wafers.


It’s not about whether or not Nvidia will be able to sell hardware to these providers - it’s about literally killing companies they are financially invested in.

Why would you invest money in a company, and then enter the market to compete with them?


Meta is awfully close.

lol! Good one...

Went from years behind to months pretty quick.

[flagged]


why so much negativity and certainty?

They have a lot of moat, i'm not sure what youa re talking about. Only amatures are using Qwen, open source stuff that is 3-8 weeks behind. Plus OpenAI has some verticals that keep people in there.

In what way do they have a moat? A cursory look at https://artificialanalysis.ai/models/gpt-6-astra#intelligenc... it lands at 61, only a single point above glm 5.3 while costing significantly more.

The only moat they appear to have is by hoarding compute, and the current trajectory of hardware shows that isn't permanent either for very long


I wish people could see how some of this reads. You are an “amateur” using a model 6-8 weeks behind? Really? Sigh.

The only good take here.

Especially since they still serve Codex-Spark, which is dogshit.

But Bonsai is garbage.

As if our EU leaders aren't a complete joke as well. Pushing ChatControl, fascism and gambling everywhere. We're just as much of a joke.


This will degrade performance significantly. LLama.cpp has had this for a while and it tanks benchmark performance. I ran GPQA on GLM 5.2 using the llama implementation and it came back 19 points under the regular results.


Or you can use parlor to chat with it directly https://github.com/fikrikarim/parlor/


Better than Opus 4.8 on complex tasks but tends to overthink. It found a bunch of bugs and architecture issues that only 5.6 Sol Max and Fable on my C++ projects.


It depends on what you need, but Krea/Klein9b/Ideogram4/Z-Image are among the best right now for text2image and Qwen Edit and Klein are probably still the best at editing.


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: