The thinking people who would find this interesting and read this are probably more than capable of understanding this and critical enough to expect that. Conversation over the title is distraction of what's important. Just stick to keeping original source title and let people vote and down vote if they don't like it. That's what votes are for.
It's regularly the case that they are simply too long to fit, so editorializing is necessary. Either cutting them slightly short, removing descriptors, excess verbs etc, is better than chopping a word in half.
Unfortunately, advertisers are getting smarter and using bots to praise their own products on Reddit. Thanks to training on genuine comments, some models are very good at sounding like a human commenter, and can easily generate a comment history with diverse interests to appear human, making them basically undetectable. So it seems like this method will, at some point, not identify the company with the best knife, but the one with the most ad spend on bot comments.
The other day I wanted to gather Reddit comments about a solar panel vendor. Claude doesn't have access to I had Gemini do some "deep research". When I fed the verbose report back to Claude it basically said it was a bunch of "hallucinated bullshit".
I've never had a Gemini Deep Research report that didn't sound like a load of pseudo-intellectual BS. It always starts with a long grandiose preamble and then sounds way too academic, almost like a caricature of academia.
I found it much more useful to go to a knife shop and handle a whole bunch of knives for myself. They’re all pretty similar besides material, so not much signal you’re going to be able to glean from people arguing on reddit.
That's called taking responsibility for your actions and being an informed consumer or a critical thinker.
Unfortunately that is completely undoable for most people nowadays. That's seen as "too much work" for most people. That's something that people see as a waste of time and should be done by something else for them. A knife should be a 5 second one-click buy, then when it sucks people will complain that all their options are bad.
You spent time/money traveling to a shop? In 2026!? Omg what a waste of time! Why didn't you aggregate the 500 product review sites and Reddit comments to pick 5 possible knives? (Sarcasm)
Most of the attributes don’t matter. Most people would be much better off with a $50 Victorinox that they kept sharp and a wood cutting board they maintained than upgrading the knife. If you are using it all day there are definitely looking things from a comfort perspective but for most homes, does not matter.
I think this is the way. An LLM is an expensive general purpose tool and for repeatable tasks, after it's clarified the process flow, it builds cheaper special purpose tools for each step
100%. We now have super general tools that reduce the cost to build other specific tools. My favorite thing with LLMs has been building a ton of little utilities for work and personal that I could've built before, but never had the time at work or the want to spend time on in my personal time.
Isn't it just learning to map specific words, from the "knife world", to the correct class? If so, a simple dictionary would fit.
What I think is a better way to validate is to split train/validation by words used presented in NER classes (like, it should be able to find new brands never seen before). It is a interesting problem.
That's a key piece of the article. He 'trusts' Gemini to classify the posts, and never hand validates anything.
And he continues building on top of that shakey trust. Is this a good idea?
It just depends on what you're doing: sounds like he's just having fun with a hobby, so it's harmless.
All he's really done is work more efficiently to save token cost and time. But he hasn't validated anything, so there's no telling if he's wasting time or not. It's a hobby though....
> So I scrape the Reddit threads where people argue about them and pull out every brand, model and steel they mention, to see what is getting bought and argued about.
Funny, normal harness with web search would one shot this, after like 20 minute search. "Replacing gemini" usually means instaling some(any)thing else, not digging deeper to get out of hole called gemini.
As for knifes, it is all same. Just do not buy total junk. Japanese knifes are way way overpriced.
This article was written with the assistance of AI. If that bothers you, stop reading here. The numbers are real: every score comes from the ten training runs described below, and the full run log is in the linked knife.day write-up.
"I didn't write any of this, but you should still trust that the remaining work, ideas and observations are all mine."
To include a disclaimer like this is to fail to recognize that "real numbers" are way less meaningful when there's clear evidence that the prompter of the LLM is not really qualified to validate them.
Request subtopic be changed to “I used Gemini to design a tool to replace specific uses of Gemini.”
The thinking people who would find this interesting and read this are probably more than capable of understanding this and critical enough to expect that. Conversation over the title is distraction of what's important. Just stick to keeping original source title and let people vote and down vote if they don't like it. That's what votes are for.
Or just: "I used Gemini to generate a tool".
titles on HN should match the original, rare exceptions
Use the original title: Submit the actual title from the source page unless it is misleading or linkbait.
And this can be said to be misleading in my opinion
I agree, this is misleading.
If he wanted a Gemini replacement verbatim, its called locally inferring it's sibling, Gemma.
It's regularly the case that they are simply too long to fit, so editorializing is necessary. Either cutting them slightly short, removing descriptors, excess verbs etc, is better than chopping a word in half.
[dead]
Unfortunately, advertisers are getting smarter and using bots to praise their own products on Reddit. Thanks to training on genuine comments, some models are very good at sounding like a human commenter, and can easily generate a comment history with diverse interests to appear human, making them basically undetectable. So it seems like this method will, at some point, not identify the company with the best knife, but the one with the most ad spend on bot comments.
who the fuck is using reddit for product recommendations? reddit was astroturfed in like 2015.
If not earlier. Reddit has been highly manipulated since at least 2015.
I used to buy/sell reddit usernames for astroturfing and it was always hilarious seeing one of my usernames on the front page.
The other day I wanted to gather Reddit comments about a solar panel vendor. Claude doesn't have access to I had Gemini do some "deep research". When I fed the verbose report back to Claude it basically said it was a bunch of "hallucinated bullshit".
Use Claude code and the chrome plugin to access reddit.
I've never had a Gemini Deep Research report that didn't sound like a load of pseudo-intellectual BS. It always starts with a long grandiose preamble and then sounds way too academic, almost like a caricature of academia.
Well was it hallucinated bs?
[dead]
It’s unreadable but then again it says it on top, but it really is so why post it
I was hoping he tricked Gemini into running the training on the cluster that Gemini itself is running on. That would be novel!
I found it much more useful to go to a knife shop and handle a whole bunch of knives for myself. They’re all pretty similar besides material, so not much signal you’re going to be able to glean from people arguing on reddit.
That's called taking responsibility for your actions and being an informed consumer or a critical thinker.
Unfortunately that is completely undoable for most people nowadays. That's seen as "too much work" for most people. That's something that people see as a waste of time and should be done by something else for them. A knife should be a 5 second one-click buy, then when it sucks people will complain that all their options are bad.
You spent time/money traveling to a shop? In 2026!? Omg what a waste of time! Why didn't you aggregate the 500 product review sites and Reddit comments to pick 5 possible knives? (Sarcasm)
Most of the attributes don’t matter. Most people would be much better off with a $50 Victorinox that they kept sharp and a wood cutting board they maintained than upgrading the knife. If you are using it all day there are definitely looking things from a comfort perspective but for most homes, does not matter.
I think this is the way. An LLM is an expensive general purpose tool and for repeatable tasks, after it's clarified the process flow, it builds cheaper special purpose tools for each step
100%. We now have super general tools that reduce the cost to build other specific tools. My favorite thing with LLMs has been building a ton of little utilities for work and personal that I could've built before, but never had the time at work or the want to spend time on in my personal time.
Isn't it just learning to map specific words, from the "knife world", to the correct class? If so, a simple dictionary would fit. What I think is a better way to validate is to split train/validation by words used presented in NER classes (like, it should be able to find new brands never seen before). It is a interesting problem.
That's a key piece of the article. He 'trusts' Gemini to classify the posts, and never hand validates anything.
And he continues building on top of that shakey trust. Is this a good idea?
It just depends on what you're doing: sounds like he's just having fun with a hobby, so it's harmless.
All he's really done is work more efficiently to save token cost and time. But he hasn't validated anything, so there's no telling if he's wasting time or not. It's a hobby though....
It wasn’t as simple as just mapping brands… for one I don’t have a complete dictionary of all the brands that could be mention.
Other problems are:
Brand names that match common English words, brand names that are also item models used by another brand… etc
> after 24 minutes on a Tesla T4 […] about $2.50 of GPU time
Who's the scam cloud provider who sells T4 GPUs for $6.25/hour?! That's B300 territory!
I think the AI writing disclaimer was a decent touch.
> This article was written with the assistance of AI. If that bothers you, stop reading here.
Okay!
I feel like this warning solves it. No shame or trouble needed for anyone.
I stopped reading three sentences in when I realized it was getting hard to follow. AI explains it.
[dead]
> This article was written with the assistance of AI. If that bothers you, stop reading here.
Genuinely appreciate the honesty. If you believe there’s nothing wrong about writing with AI, there’s no reason to not own up to it.
> So I scrape the Reddit threads where people argue about them and pull out every brand, model and steel they mention, to see what is getting bought and argued about.
Funny, normal harness with web search would one shot this, after like 20 minute search. "Replacing gemini" usually means instaling some(any)thing else, not digging deeper to get out of hole called gemini.
As for knifes, it is all same. Just do not buy total junk. Japanese knifes are way way overpriced.
I wish they told us which model wrote the article so i can be sure to avoid it in the future. This is brutal to try and read
what do you do when a new brand of knife comes out?
To include a disclaimer like this is to fail to recognize that "real numbers" are way less meaningful when there's clear evidence that the prompter of the LLM is not really qualified to validate them.
The reader assigns value to what is read. I found some value in what I read, even with all the holes.
It's a different way to think about a problem, even if it's not applicable in all areas.
I had Gemini read this article and write its summary to /dev/null
The more intelligent AI become, the less moat it has
I'm sorry Dave, I can't do that, unless you upgrade to a premium enterprise subscription.
post training dataset for GLiNER is pretty small though.
Thank you to the author for disclosing slop writing up front. I appreciate you respecting your readers time.
[dead]
[dead]
[dead]
[dead]