Hacker Newsnew | past | comments | ask | show | jobs | submit | xlayn's commentslogin

unleashes unparallel, unrivaled, totally new, amazing, exciting new ways of interacting with the MOST POW3RFUL!!!!!

then all the photos either what's on the inside it's repeated whats on the outside,

or... and hold your breath... you can show on one side the calendar and on the other side the clock... now you can breath...

Meh, another 2 or 3$k gadget that does the same as the previous one?


The more advanced the effort to detect human vs AI the more advanced the AI will be to write just like human. I wonder if at some point it will become totally human like, and then what will be the reaction? will you accept it? will it be tuned and targetted just like ads to you so it can be eloquent and influence you?


No they are not, try to change jobs and if you can't or takes a really long time you know this is some of the finest baloney. And no, it's not creating a lot of jobs, one one of the kind that will be destroyed by AI no, and that's the point... you built something that removes the pricey laywer, software developer, insurance policy writer, designer, etc... and replace it with a ONE TIME go work under the sun with no ac in the middle of nowhere as a contractor of the contractor of the contractor... So sure, I'm a dooms-scroll-lover, sure the 10000000 new part time possitions with no benefit of any kind will be the biggest work boom EVER... you will just go to compete with all the other people already doing that... and you know what happens when there is a queue of 100 to flip burgers or "highly technical non AI replaceable work like weld or electricist" right??

"you built something that removes the pricey laywer, software developer, insurance policy writer, designer"

This statement is untrue. The people have the judgement, the AI does the rote work. The rest is marketing fluff and/or leadership fumbles. You can't get rid of judgement (unless of course leadership is fumbling).

With more AI, you got more radiology requests, and more radiologists [1] and better health results due to more abundant measurement. There is no evidence that the same isn't happening to the rest of the professions. And to support it all, we have more infrastructure construction (datacenters).

This is the textbook basics of a growing economy with growing productivity, even though it feels paradoxical [2]

[1]: https://www.ft.com/content/f2e03bd9-af67-45c4-8e1e-79978b5bc...

[2]: https://en.wikipedia.org/wiki/Jevons_paradox


Jevons Paradox doesn't say more radiologist will have more work and will be paid at the same price, it says that there will be more work for radiologists because now they are paid less and because of that people is able to afford more radiologists.

"This statement is untrue. The people have the judgement, the AI does the rote work. The rest is marketing fluff and/or leadership fumbles. You can't get rid of judgement (unless of course leadership is fumbling)."

Let me refine it more... now the same lawyer/marketing/software engineer can work just approving the work of AI, reviewing it's correct so the exact same lawyer/marketing/software engineer can now perform X more times the amount of work they previously did, IF we don't triple the amount of work, then the rules of "free market" means that offer now exceeds by X the demand and you can now depress wages as much as you want because you can fire (x-1)/x of the people doing the work and the pool of people willing or needing to do that is increased.

This does not contradict at all what I stated above.

The ultimate proof of this is how much the "calculator people" now earns now that calculators were created... and for reference this old ad about the effect of calculators on the number of engineers

[0]: https://www.globalnerdy.com/wordpress/wp-content/uploads/200...


"radiologists because now they are paid less and because of that people is able to afford more radiologists."

But again, the opposite is happening. Radiologist salaries are up 9% YoY [1]. And you can now get your radiology done more cheaply [2]. Literal win-win.

Radiologists are not human calculators, neither are lawyers, accountants or insurance adjusters. As we learned very quickly, their job is judgement and human relationships and you can't cut that out

You did correctly lay out that if you can't grow your business by increasing demand, you are shrinking. And yes shrinking leads to layoffs regardless of AI.

Very cliched now but: You are not being replaced by AI, you are being replaced by someone who's using AI, who is more productive with it. But specifically, this is true for individuals and company-wide. Less productive organizations will be pushed out if they can't lower their prices and meet the increasing demand, and the only way they do that for modern white-collar work (not rote 1950's calculator churning) is to hire more AI-ready staff.

FWIW, The IBM ad is literal marketing fluff. More engineers were employed in the immediate aftermath of the mainframe than anytime before.

EDIT: I'm not trying to argue with you, I'm trying to show you things most headline news refuses to show because it's boring and counter-narrative. Shoutout to the economist.

[1]: https://fortune.com/article/ai-godfather-radiologists-obsole...

[2]: https://pmc.ncbi.nlm.nih.gov/articles/PMC10986240/


Exactly spot on.

Why are people so stupid?

E.g. take the radiology example. In the past people would not have been sent to a radiologist due to resource constraints.

But because of AI you now more capacity - therefore there is an INCREASE in radiology requests.

And this is going to be the common theme - there are lots of things that people are doing right now has an implicit constraint. AI will unlock new capacity so we will see MORE of it -> this will actually create a bizarre scenario of a labour shortage in many respects.

I know that's not what Amodei and Altman want to hear but... ha.

THis is why people are so bad at predictions - they look at one variable move and forget the rest / many 'unknowns' they did not foresee.


if you're going by anecdata, myself and a few of my colleagues (senior ics) switched jobs relatively easily for a solid pay increase within the last year. definitely easier than post-covid

engineering (hardware, software) and data center construction are going through a boom cycle right now and will eventually bust, and so on and so forth


AI has completely broken interviewing as well. The asks I am getting in the few interviews I can achieve are schizophrenic.

They have pretty uniformly wanted me to demonstrate using AI but then freak out and interrogate me for every prompt and then want me to demonstrate I can do the task by hand. And that’s only after getting through the AI run filtering that gives a different score everytime your unchanged resume is run through the model.

It’s dark times for employment currently unless you’ve got a rare nepotism path.


I don’t know if those companies are doing it for the right reasons but I’ve implemented a policy of showing me that you can do the job manually as well.

The reason being that if you can’t then you probably can’t tell when the AI is wrong.

My company is strictly Cybersecurity/IR though so that may vary with other specialties.


If the interview was specifically about doing it manually that’s one thing.

I called them schizophrenic because the companies I’ve interviewed with seem like they desperately want to be on the AI bandwagon and want you to demonstrate the ability to effectively use AI, but they don’t currently know how to evaluate that and panic and default back to their old heuristics.

It is very difficult in 45-50 minutes to complete a task when the criteria change partway through.


Oh yeah no that’s just a terrible interview process.

I think schizophrenic is a great way to describe that.


It was always better to find job via recommendations. Both for companies and candidates. Old leetcode style interviews were no better at identifying good candidates for given job.

> It’s dark times for employment currently unless you’ve got a rare nepotism path.

In the civilized society, we call that "recommendation".


And the robots are just a few years away. I thought that spatial reasoning was going to be a lot harder for AI, but Astra is showing that progress on that front is also moving very rapidly.

humanoid robots are decades away, and they are not going to be good or cheap for a long time.

> humanoid robots are decades away

I think the world has gone too crazy to make any such claim. They could be decades away, they could be a year away.

Remember that it doesn't have to be necessarily better than a human, it just needs to make economical sense. A robot that fails 20% of the time but costs 0.10$/h to run beats a person for many tasks.


"humanoid" robots are as relevant to commercial robotics as AGI to LLM, that is, not at all. As long as they can move on wheels and operate 2 hands (doesn't matter how many fingers), the application is huge already.

Decades? Or less than a decade?

Less than a decade with 99.999% certainty.


Hey carloslfu, kudos from the other side of the internet, don't get down on people nitpicking everything here, experimenting and discovering is part of learning so keep going!, remember this is the place that said dropbox was dumb and could be replaced by a script.

I pay a $100 subscription which feels infinite for me, even using only fable... 2 weeks ago I started getting notifications of running out of credit... to me it feels more like 30% than 17%... OHHHH and surprise, I'm in my 50% boosted/higher limits... so this means I'll probably hit the $100 limit now constantly... I use the thing a bit... yesterday and it shows 14% of the week allocation used...

I worried about this myself. 50% higher limits were nice but the flip side is you get used to it and then it just feels like they're taking 50% away. I'm not sure it's going to have the effect that Anthropic was expecting....or maybe that was the point. Who knows. I can see both perspectives.

For the impatient, I merged llama.cpp tentative branches to get it running here https://github.com/alainnothere/llama.cpp/tree/disk-cache-ev..., thing runs at 23.54 token/sec and my setup runs at high 30 the 3.8 dense 27B.

and this is the pelican from the iq4_xs model https://github.com/alainnothere/llama.cpp/blob/disk-cache-ev...


If you've got a DGX Spark try my little engine: https://github.com/rdaum/eider/

NVFP4 quant


This is the second post in the vibe of "you should not get angry at work, being angry is bad, and you want to be a professional bla bla bla", and what I see is some effort to start putting the blame into the worker of working conditions that make you angry.

It's like the petro-campaign of "drive as much as you want, burn fuel non stop but if you recycle then everything is good...." then you put paper in one bin, and "the other garbage" in the other... two different trucks pick it and ends up in the same landfill (google "recycle end up same landfill apple tag" and pick one)... and by the way, the whole recycle thing was a diversion to keep your focus away from the actual issue... google "recycle was a scam by petro" and pick your poison

like when you start reading... we should start getting used to earth warming... how to tolerate more heat... and yes just google


In animatrix there is this part of the video where the devised path forward is "the destruction of the sky" and you can see all this millitary and business people clapping... then the image turns to all that people as skeletons clapping...

let's burn stuff until the heat is unbearable, what can go wrong? let's mess with nature, what can go wrong...? let's destroy the clouds so we can improve the ROI in our solar farm, fuck the shade, the water and the rain...


That part of Animatrix never made sense - like if the rogue machines could not switch to nuclear power or similar.

Which they did + bullshited humans abou using them as a power source, while just doing that for fun.


I do use 2 amd gpus and I get high 40 for generation, 500 for pp and low 20/100 by the end of the context of 256k.

llama-server --host 0.0.0.0 --port 8089 -m Qwen3.8-27B-UD-Q8_u.gguf --spec-type draft-mtp,ngram-mod --spec-draft-n-max 3 --spec-draft-n-min 1

if you have an igpu and want to exclude or just use some gpus you can use

--device Vulkan3,Vulkan2,Vulkan1

in my case vulkan because of amd, you can see your devices with

llama-server2 --list-devices

Available devices: Vulkan0: AMD Radeon Graphics (RADV RAPHAEL_MENDOCINO) (33515 MiB, 29349 MiB free) Vulkan1: AMD Radeon RX 7900 XTX (RADV NAVI31) (24560 MiB, 4911 MiB free) Vulkan2: AMD Radeon RX 7900 XTX (RADV NAVI31) (24560 MiB, 7681 MiB free)


Hey Unsloth, your gguf are the first ones I look for when I want to download a gguf model. Today I was trying in fact to see, what's the smallest Qwen3.8-27B that I could run and get good results, say restricting it to 16GB of ram.. so I went, pick up the Qwen3.8-27B-UD-IQ2_XXS.gguf and them BAM, error on MTP... now I understand why after reading your announcement. Beyond the space saving, why removing the MTP? improves speed exactly for the group that could benefit from it.


The reason for running those insanely low quants is to fit in extremely limited memory budgets. The first thing you sacrifice is speed, then context and accuracy (up to you in which order). IQ2_XXS and below is desperate/proof of concept territory. If you have a spare half gig for the MTP drafter, run a larger quant instead, it will be less incoherent, and damn the speed, it won't be garbage at least. Only around Q4 I'd allocate the comparative luxury of more memory for a speed increase. At least on a dense model. MTP makes a lot more sense (but helps statistically a bit less) on an MoE.

Qwaiting for that 3.8-35B-A3B


Hey we did not remove the MTP for sizes above 8GiB - but yes for small GGUFs under 8 ish GiB, we removed the MTP module (IQ2_XXS and lower), because it's 500MiB to 750MiB in size, and on small 8 GiB machines, even 500MiB is needed.

As someone in the comments said we made a separate Q4_0 MTP if that's helpful so you can use that.

But I would suggest using UD-IQ3_XXS for 10.9GB for 16GB machines or Q2_K_XL


Daniel, question I got the Qwen3.8-27B-UD-Q2_K_XL.gguf from https://huggingface.co/unsloth/Qwen3.8-27B-GGUF?show_file_in... and continue with my testing, but the model quickly felt into a loop of asking the same thing over and over again, I have seen the MOE do that but not the dense ones.

And I had similar experiences when Qwen3.8-27B unsloth images just came out with the full Q8_K_XL, I'm using an AMD setup which has modifications to save to disk the kv, but your (assuming you are part of the unsloth team) for some reason have been giving me similar issues.

I tried https://huggingface.co/mradermacher/Qwen3.8-27B-Uncensored-G... the 8 bit, 6 and 2 bit... the 2 bit almost use the complete KV doing it's thing and didn't loop itself.

It can be something in my setup, there is a very high chance of that, but the previous 3.6 images from qwen, the 27B, the 31A3 and 122 they are all unsloth and did work on my setup without issues...

Again could be my setup... let me know if there is any data I can supply to you to debug if needed.


Are you using the recommended settings for temperature and such? https://unsloth.ai/docs/models/qwen3.8#recommended-settings

Often times I run into issues like this it’s because I am using settings for a different model or just forget to set them up.


>But I would suggest using UD-IQ3_XXS for 10.9GB for 16GB machines or Q2_K_XL

For those of us with a 16GB GPU, how do they compare with ExllamaV4 at 4-bit (4.0bpw)?

It looks like that fits in 12.5GB of VRAM since embedding are left in DRAM, Unsloth Studio and other llama.cpp derivatives have to load these weights in VRAM for tied embedding models like Qwen3.8.

ExllamaV3 4.0bpw fits in 12.5G of VRAM and beats IQ4_XS according to the measurements here: [turboderp/Qwen3.8-27B-exl3](https://huggingface.co/turboderp/Qwen3.8-27B-exl3)

But those were compared against UD2.0 I guess. Also plans to support these (SOTA) quants in Unsloth Studio?


you can still have it, no?

> We also removed the MTP module from smaller quants under UD-Q2_K_XL (8.37GB and lower) to converse around 500MB of disk space - you can use the Q4_0 MTP separate module if needed


my bad, you are totally right, thanks!


Q2 quantization is basically giving a capable model a lobotomy. It will not accurately represent how smart or capable something like qwen 3.8 27B in Q8 will be.


Sure, but this is true for all lossy compression (audio, images, etc.)

Given 16GB of VRAM, what will give me the best experience in OpenCode? Currently using Qwen3.8_Q_3


If it's a Nvidia card 3000 series or newer, I'd try 4.0bpw ExllamaV3 if you haven't already. Otherwise it look like UD3.0 Q3_K_XL based on the Unsloth blog post.


I'll try it out, thanks. (Using an AMD 9070 XT)


probably the best experience would be deepseek v4 flash 0731 (it takes about 170GB RAM on the server side for the full thing and RAM reserved for 1M context) via opencode's $10 a month plan until you use that up, it's either Q8 or full precision. Assuming you're ok with doing things with external inference.


Why would they have listed how much vram they had if they were looking to rent gpu time on someone else's machine?


A casual review of my comment history would show that I've been nothing but the biggest proponent of running models locally, and I do so myself a great deal. But one also has to be realistic about the capabilities of what you can do in a 16GB GPU these days. I already said an extra small Q2 quantization was effectively lobotomized so I didn't want to repeat myself.

This person has basically run into the limit of state of the art for even a modestly sized local model (this isn't deepseek v4 flash 0731 Q8 which I am running myself locally on a great deal more hardware), this is a 27B dense, but they're just not going to have a good time if they expect good quality results out of a Q2. The choices are either upgrade hardware or pay for external inference.


Fine, but they already said they are using a Q3, so Q2 being unusable (disagree, but whatever) isn’t helpful new info.


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: