- cross-posted to:
- [email protected]
- cross-posted to:
- [email protected]
So typing isn’t fast enough to burn all the daily tokens?
The faster you burn them the sooner you knock off for the day.
the faster you burn them the sooner you burnout
Automate it
Have all chatbots you have access to talk to each other through each other to try to solve the problem
“knock off for the day?” Not until you’ve trained your replacement (LLM)! Then you can take all the time you like, income-free
If I’m being forced to spend AI tokens up in a day I’m spending the rest job hunting.
if you talk AND type you can burn twice the tokens and get the same result! it’s a win win win
Have you seen people type? Toich-typing is a lost artform.
They don’t even have to do that. Just don’t pack them in like sardines.
Now that’s a waste if i have seen one.
Most type faster than they speak, no?
No.
You’d think so, given how most modern vibe coders talk. But no, they’re so bad at typing, they can’t do either!
Techbros will have workers do anything but work from home.
Or have single-person offices instead of an open space
You will love the Open Space Farm.
You mean cubicles?
Now I gotta rewatch Severance!No!
in this one picture, I see generations of so many individual great ideas coalescing into one fantastically bad idea (and stupidly comical consequence) that is so… bizarre that I can both simultaneously understand why nobody really saw it coming, and am in low-key disbelief that this is even real…
I mean, that’s a steno mask, and anyone who’s had issues with hand pain but wants to communicate via text has probably wished for something resembling this. The problem is that they’re obnoxiously expensive. (And they look ridiculous, but that’s its own issue.)
When I’m in public, I wish people had these. I neither need nor want to listen to your phone conversations.
Their price is so dumb, that thing shouldn’t cost that much
Just like those accessibility tools that let you use control the mouse pointer using only your mouth etc etc, they have stupid prices when some aren’t even that complex
They’re a niche product that has a very solid use case (court transcripts). They require far less training than a stenotype, but because they’re so niche, the people who do make them can more-or-less charge what they want. If we’re being less skeptical, the low demand means the manufacturers can’t take advantage of the efficiencies of scale, so the price remains high.
of course it’s not, because every laptop and airpod has
noise cancellinginput isolationBut as far as rage bait goes, top tier
I think this is a microphone with noise isolation, rather than noise cancelation from speakers.
Thanks, edited for meaning
Isn’t it cheaper and quieter to just type out your prompts?
This is akin to people who have conversations on speakerphone in public places.
you assume that vibe-coders can actually touch type. or type at all.
I assumed they were all working from their iphones
Or be capable of critical thinking.
AI needs to measure the level of confidence in your voice, to calibrate its bullshit accordingly
I wonder if anyone owns a patent on this idea.
Speaking is faster than typing, I guess?
Maybe for Eminem.
Older people such as myself tend to hate voice-to-text I think because it was so awful in the past. And if you screwed up with it in the past it was a less understandable excuse that “I was using voice-to-text.” And because we were all forced in some way to learn to type well.
Voice to text works a little bit better now. And I think younger people know everyone else uses it and to forgive when it screws up.
I hate it because it’s generally intensely US-centric. Not understanding non-US accents or terms.
It worked better five years ago before they replaced the existing algorithms with AI bullshit. Keeps adding slurs to my dictionary too since they replaced keyboard prediction with modern AI and so many people use slurs Google just assumes I do too.
apparently audio and images are more efficient compared to text for multimodal models?
I’m going to need significant levels of convincing. Computers have always preferred specificity and accuracy, it’s half the reason I’m in my current position (MSP Escalations/level 3, half of my success at fixing issues is being extremely specific in looking up exact error messages instead of paraphrasing).
This isn’t a defense of AI; on the contrary, it’s my doubt that AI can read intentions/inflection/emotion better than just writing out what you actually want.
Deepseek recently published a paper in which they describe that vision tokens contain more information than text tokens and that this can be used to compress context.
We present DeepSeek-OCR as an initial investigation into the feasibility of compressing long contexts via optical 2D mapping.
Experiments show that when the number of text tokens is within 10 times that of vision tokens (i.e., a compression ratio < 10×), the model can achieve decoding (OCR) precision of 97%. Even at a compression ratio of 20×, the OCR accuracy still remains at about 60%. This shows considerable promise for research areas such as historical long-context compression and memory forgetting mechanisms in LLMs.
It reminds me of LLM caveman speak, it used to have another option to use Chinese instead of English. A language like Chinese is seemingly better at encoding information in fewer tokens and I think this is the same mechanism why OCR tokens work so well.
That said, I also doubt that voice messages are more efficient than text prompts, but it’s best not to waste too much time engaging with these sorts of LinkedIn posts (and LinkedIn in general).
LLMs don’t need accuracy. This just boils down to speaking being faster than typing, especially if your thought isn’t fully formulated.
deleted by creator
As far as I know, these workflows typically involve a transcription model to convert the audio to text, and then passing the text to the model.
I would rather do almost anything than talk to a device, except in very specific circumstances.
I set timers and play music on a smart speaker somewhat often.
Occasionally, when I am alone, know exactly what I want to say, and my hands are full, I might dictate a text message.
But other than that, I will not be talking to my device, thanks. The human voice is primarily for talking to other humans, with all the imprecision and uncertainty and emotional resonance that entails. Keyboards are great tools designed for precise computer input, and I would like to continue to use them.
i dont think ive ever used a speach computer interface that wasnt hot garbage and misunderstood half of what i said unless i talked to it really slow like it was an idiot. pretty much every time it would have been faster to use the tactile interface. i dont even have that much of an accent compared to generic american.
Hook it up to a bong and you’ve suddenly made work a hell of a lot more interesting.
I was gonna fill mine with oatmeal and chow down, but I think there’s some good work we can do putting both our minds together on this one.
Oatbong?
The real vibe coding
Or, listen me out, they could work from home.
Working from home doesn’t appeal to the emotional needs of fragile managers.
It doesn’t appeal to the emotional needs of a bunch of my colleagues who are in the office every day voluntarily either.
Yeah, I have colleagues who choose to work in the office when work from home is available because they like the separation of work from home, don’t have a good spot to work from home, are aware they would be distracted at home, prefer to see other people in person, and a bunch of other reasons. At least they get a choice!
Every day I’m a little surprised there’s no news story of some workers beating their “no, you have to come into the office” manager to death. They’ve got means, motive, and opportunity, and it’s extra funny because if they’d been allowed to work at home they wouldn’t have at least two of those.
But really we’re ruled by the worst of us. Cowards and fools.
Maybe unionizing is safer than hitting the decision makers with an office chair while screaming “you made this possible” until they can’t even cry anymore.
Well what I’m trying to tell you is that there are probably more people than you realise who want to be in the office. My partner, and a bunch of my coworkers, hated being forced to work from home during the pandemic. So maybe that’s part of the reason.
Oh, I read your thing backwards then.
I can’t imagine wanting to go into the office on the regular. The commute. The lost time (can math out to like a 20% pay cut, if you spend two hours a day traveling + getting ready). The sickness. The lack of control over environment (temperature, sound).
Can’t relate to it. And I’m a very social person that likes interacting with people.
I live 15 minutes from work. If I got out of bed early enough I could easily bike there (and have done so).
Really the only thing I miss from work from home is the ability to take a short break once per hour to do some bodyweight exercises or kettbell swings during lunch.
But on the flipside, it’s hard for me to stay concentrated while at home. I personally get way more work done in the office.
But on the flipside, it’s hard for me to stay concentrated while at home.
While I believe that is true for you, I don’t believe it justifies the lost time, health, and environmental damage of mandating in-office for everyone.
The office is super distracting for many people.
A noise cancelling microphone isn’t a terrible invention though.
I’ve seen those things 10 years ago when somebody was helping a deaf student in university. Used one of those things to talk to a text to speech engine during the lecture. Seemed to work quite well.
I’ve also seen them on court stenographers and thought they looked cool. Not everyone needs every invention, though.
Thats true this is just the absolute worst use case scenario i can think of
They suck, you can’t get air into them without destroying the noise cancelling part so you’ve got to constantly be taking it off to take a breath / running out of oxygen mid sentence
You talk with your nose?
You don’t have allergies or colds?
Why stop at feeding? Where’s the penis and vagina tubes? We might as well be jorkin’ it too if all the rich and powerful are.
It does mention “all the other tubes”. I wonder if it will be THX 1138 style.
I hope its like a Geigeresque device that makes us sit shrimp style and cups backwards from our gooches to our mouths.
Alexa, play Spankfest 8.
Anyone needs handmade wooden furniture?
Because that’s what i’m going to be switching careers to if that trend comes to my place.
Woodworking is surprisingly popular among tech folk. It seems some hobbies just click better for techies. Bouldering is another example.
Sewing/designing clothes really clicked for me. Haven’t tried woodworking, but I imagine it scratches the same itch and utilizes similar skills: 90% of sewing is just planning, calculation, and measuring. Then, watch everything just fit together into place
Exactly like programming! /s
Never hesitate to rescue wooden pallets, but don’t waste gas on retrieving them either. You can use an iron railroad tie as a chisel for dismantling them

This looks uncomfortable and humiliating. Now if they were to make it in the form of a suppository…
If you’re brave enough, anything can become a suppository.
Did the CEOs unplug their keyboards? Wtf even is this?
Well the mask is a steno mask
Theoretically most people will speak faster than they type. You have to type around 180-200 wpm to be faster than speaking.
(I say theoretically, because usually typing speed ratings also ding you for errors, and uh, speech transcription isn’t really there, either.)
The hard part of both speech and typing is thinking about what you say. Typing nor speaking are going to change the speed I can get information into the computer.
Maybe we could ask the AI to do that thinking bit then tell us what to say.
Hey Siri, tell ChatGPT what it wants to hear to generate a million dollar code piece.
Without bugs.
I’m not really into it but one of the guys on late night linux podcast, generally resistant to LLMs and shiny new things, is a fan of speech to text for general computing, as in every input field in the OS should support speech. In a recent episode he said that he believes it to be the way of the future.
I remember
Dragon SpeechDragon Naturally Speaking saying the same thing in the 90’s. It’s improved, but not enough to make it useful as more than an aide for people who can’t type. I do agree, that for simple accessibility, it should be integrated into every field, but I doubt it’s ever going to take over.As others have noted, that it’s only technically true that dictation is faster than typing. In a practical sense, there’s a fair number of reasons why that’s not the case, including that usually thinking about the entry is what’s the slowest, and also the errors in both are typically what slows people down.
there’s also the problem of, for example, keeping entries confidential. You don’t want to speak your passwords where others can hear you.
I remember Dragon! And ViaVoice! I saw a presentation for ViaVoice in the late 90s and it blew my tiny mind.
It really was the future. And it’s… a bit better since then. Oh god thats like 30 years ago almost
It’s improved, but not enough to make it useful as more than an aide for people who can’t type.
I don’t think this is true.
There’s a locally hostable model called whisper that is very impressive.
My plumber uses speech to text to send text messages all day.
Late Night Linux guy says he uses it for microsoft teams quite a bit.
You’re only partially correct about input speed. If you want to dictate an email then yes you need to think about each word you want to say and the order in which to say them. Coupled with an LLM that problem is diminished because you can just kind of have a conversation with the LLM and tell it to draft an email.
You’re only partially correct about input speed. If you want to dictate an email then yes you need to think about each word you want to say and the order in which to say them. Coupled with an LLM that problem is diminished because you can just kind of have a conversation with the LLM and tell it to draft an email.
and how much of that conversation with an LLM is “No, what I want is…” because it assumed something; or just straight up hallucinated or the typo made it go off on a tangent?
As for whisper, I can find sources that are saying for American-English speakers in a not-noisy environment (aka the best case scenario,) the model has a word error rate between 2-8%. For reference, Dragon NaturallySpeaking had a WER of 3-5%. So I wouldn’t say that Whisper has made any substantial improvements, and they’re OpenAi. you can trust them if you want. I don’t think that’ll work out well in the long run, though.
I’d like to see the source that says Dragon’s WER in the 90s was 3-5%. I used Dragon in the 2000s and it just wasn’t comparable to the current state of the art.
whisper.cpp is an opensource implementation, although I’m not certain exactly how open.
when you’re providing context rather than instructions the tendency for a model to hallucinate or run off on a tangent is minimal, because the context you’re providing has it’s own cohesion.
I’d like to see the source that says Dragon’s WER in the 90s was 3-5%. I used Dragon in the 2000s and it just wasn’t comparable to the current state of the art.
https://dragon-medical-transcription.com/history_speech_recognition.html, for example. a lot of adverts and awards were given to it (admittedly awards like PC Mag that were probably paid advertising… but that’s why I went with Open AI’s assessment on whisper at 2%.) Dragon was boasting 99% accuracy after (admittedly months) of training; and it frequently reached it. there were some gotchas in that- the months-long training was a big one. The other was that you frequently had to slow down and be careful to enunciate that you don’t have to do with modern systems (including the MS versions of Dragon- they bought it out at some point)
whisper.cpp is an opensource implementation, although I’m not certain exactly how open.
It’s on the MIT license, if that helps. I take issue with anything OpenAI is involved in. for oh-so-many reasons.
If only there was some sort of device that would allow you to input those commands without interfering with others, or vice versa… But it’s probably just a dream…
Started hating working in tech a year ago and got out.
What’d you move to?
Retired and living from investments. I’m 51 so was working in tech for a long time.
Start investing if you aren’t.
Lol you think stocks will be worth shit in 20 years? This fish is going belly up and the AI IPOs are about to steal all of our retirememt funds (if you even had any).
Anyone 30 and younger knows this shit is temporary at best and collapse is within our life times.
Anyone 30 or younger has their entire brain filled with social media programming and fear. :)
The entire idea is to make you scared of saving, you dummy. I mean that in a funny way. You should never make decisions based on what social media is telling you. Social media is programming, gossip, big headlines… Building on news. Who owns the news? Billionaries. Who owns the platforms? Billionaries.
They’re not gonna steal it for long because they themselves will tank soon from all the overvaluation. They may bring about the collapse.
You know, this “soon collapse is coming” has been said since 2019.
In my opinion, the fear of collapse is exactly what they want you to feel, so you don’t make any money from the dystopian build-out.
But they are making money on it. Oh yes. Stocks go up because huge investment bankers are buying all the stocks they are telling you to be afraid of in the news.
You can sit and have no stocks and lose no money if you want, but you are betting on a crash that you are being told will happen “soon”… :) We will see. Chances are, you miss making a lot of money on this and afterwards, you will feel a bit sad you missed it.




























