Since this is for the mac you really should be using apple's vision framework for OCR. It smokes tesseract in both speed and accuracy.
Edit: I'm curious which LLM was used to generate the code. I fed the title of your post to claude/deepseek/qwen/codex asking to recommend a stack for this project, expecting to frown thinking that they still recommend tesseract. However, I found that they all recommend apple's vision framework. In fact the latest model to recommend Tesseract is gpt-4.1.
robotmay 57 minutes ago [-]
Mostly unrelated but fun thing I discovered earlier this year with Apple's vision - if you have text both correctly oriented and upside down in the same image, it likes to interpret the upside down text as a Cyrillic alphabet. I was trying to use it to read the text on camera lenses and it came up with all sorts of bizarre interpretations. If anyone's interested, I got around it by splitting the text at a point and unrolling it into a straight line before running OCR on it.
vavkamil 43 minutes ago [-]
I recently tried to recover text from 8 frames of an office-shot YouTube video, where only a small, blurry portion of a computer screen was visible. After spending half a day with Astra on it, the conclusion was that it’s not possible to read.
Later that evening, I just paused the YouTube video on my phone, circled the part of the display with Google Lens, and it read the whole thing with pretty good accuracy. It was mindblowing :)
spiderfarmer 2 hours ago [-]
So the prompt was probably something like: build me x using Tesseract.
daveguy 2 hours ago [-]
nullsanity got downvoted into oblivion, but they are correct. This is one of the many reasons why vibe coding produces worse software. The code that is generated and the best practice recommendations are completely separate. They both come from a distribution of "most common", and best practice is rarely common. Especially when a practice is first established, or in a specific niche.
thih9 1 hours ago [-]
Unless you are referencing some existing non vibe coded app, complaints the project being vibe coded may be just too generic at this point.
A hand coded electron project would have a discussion about electron vs native. Relevant in general but off topic in the context of this particular app.
xp84 1 hours ago [-]
Isn’t that why you use a plan mode? Or better, use a plan mode, revise and critique the plan, and then let it proceed to build?
baxtr 53 minutes ago [-]
Maybe the real question is: who cares?
What’s the downside risk of having "worse software" when you’re just ideating and putting things out there to see how people like it.
spiderfarmer 2 hours ago [-]
On the flipside, it also leads to better software.
allenleee 41 minutes ago [-]
[dead]
nullsanity 2 hours ago [-]
And this is why vibe coders always make inferior software.
doubleorseven 4 minutes ago [-]
how long does it takes for a FHD 90 minutes asset?
alt227 3 hours ago [-]
Slightly offtopic, but made me wonder.
Can you copyright things like this now that LLMs exist? I mean, up until now if a small startup has a great idea they will get bought out by big tech which will integrate (or kill) their tech. But now with LLMs can the likes of OpenAI just tell their model to make something that works similar to X (such as this project) and then get round copying laws and negate being behind the curve?
EDIT: switched to the correct spelling of copyright.
tough 3 hours ago [-]
copywriting is the art of writing copy for products/marketing
maybe you were thinking of sherlocking [1]
I don't see why LLMs should make the legal part different. It's like saying if you can copyright a book considering LLMs now can copy / write a new one in just one hour. Making it faster doesn't change the legal aspect / ownership of something.
modzu 39 minutes ago [-]
yes afaik copyright law has nothing to say about being sherlocked. it's happening - virtually any software you can think of has some shit slop clone out there already
allenleee 21 minutes ago [-]
[dead]
qlasisi15 54 minutes ago [-]
This would really help in video editing.
Thanks!
allenleee 21 minutes ago [-]
Glad you like it :)!
hn3ufz62f7 4 hours ago [-]
Having built something similar with CLIP on an M1, frame sampling rate is the whole ballgame. One frame a second on 12k videos is days, keyframes only got me to an overnight run.
pezgordo 3 hours ago [-]
Maybe you need a minimal downscale version as well, I heard is very common technique in the video editing world.
Based on my experience, sampling rate can be tricky if what you are looking for lasted less than interval period.
Forgeties79 2 hours ago [-]
Proxies. You transcode proxies from the original media, edit off those, then you use OM for the final render. NLE’s usually let you flip between them.
xnx 3 hours ago [-]
Have you tried scene detection? I would guess camera cuts are even less frequent than keyframes.
mistrial9 2 hours ago [-]
> 12k videos ?
you are pirating first-release movies for commercial purposes?
Hnrobert42 2 hours ago [-]
I think 12K is quantity not resolution.
measure2xcut1x 3 hours ago [-]
How well do you think this would work on stock photography on m1 mac with 32GB ram? For example I'd like to be able to search a folder of ~2k photos for houses with palm trees. Or find photos of kitchens, or find photos of desert southwest landscapes.
lucideer 3 hours ago [-]
possibly off-topic, but for anyone interested in this on a more cross-platform / holistic basis, Immich does this
(& by "this" I mean an approximate AI search for photos & videos - I can't account for the "every frame", nor for the comparative search quality)
yt1998 2 hours ago [-]
Why chose CLIP to do this. Have you tried small VLMs like Qwen-VL? I believe those models have video encoders can better perform at this scenario.
54 minutes ago [-]
stephenitis 4 hours ago [-]
I like the entire premise, the one thing stopping me from trying this is not knowing the time scales that I will need to set my computer aside for the processing of large folders of video frames, or my photos library's videos, some 12,000 videos
qprofyeh 3 hours ago [-]
This is cool. Any way to search for People / faces / pets? Like on iOS?
aavisangle 4 hours ago [-]
How do you do that? What's the architecture? Can you guide on that?
Jeeetendra 1 hours ago [-]
jumping straight to the right moment in a video is the useful bit here. how often does the balanced sampling miss something that only appears for a second or two?
freecodeio 3 hours ago [-]
would be lovely if picture embeddings were attached to the file by the camera but one can only dream of such futures
collingreen 2 hours ago [-]
Would require you to be locked in to the one embedding model in the camera though and cameras would need to use the same or be incompatible. Would be fine with a standard model like CLIP but would leave a lot of potential on the table compared to a good way to do your own embedding for everything.
cpursley 2 hours ago [-]
Why is this a JS bloatware instead of native or Rust which is easier than ever now with LLM coding tools.
collingreen 2 hours ago [-]
Because you haven't rereleased your own fork in rust yet! The "LLMs can do it" cuts both ways. If that doesn't sound worth your time then it's silly to rudely suggest it is worth someone else's.
cpursley 56 minutes ago [-]
Fair enough, but seriously, my machine has taken a beating by all these JavaScript and electron thingies.
mannanj 28 minutes ago [-]
its a cool project, but I dont want to consume someones ai slop to discern whats true. if the human wrote the page in their own words, I would have considered using it.
otherwise, I can just make my own with my own ai. why consume someones slop when I can eat my own.
Rendered at 16:22:49 GMT+0000 (Coordinated Universal Time) with Vercel.
Edit: I'm curious which LLM was used to generate the code. I fed the title of your post to claude/deepseek/qwen/codex asking to recommend a stack for this project, expecting to frown thinking that they still recommend tesseract. However, I found that they all recommend apple's vision framework. In fact the latest model to recommend Tesseract is gpt-4.1.
Later that evening, I just paused the YouTube video on my phone, circled the part of the display with Google Lens, and it read the whole thing with pretty good accuracy. It was mindblowing :)
A hand coded electron project would have a discussion about electron vs native. Relevant in general but off topic in the context of this particular app.
What’s the downside risk of having "worse software" when you’re just ideating and putting things out there to see how people like it.
Can you copyright things like this now that LLMs exist? I mean, up until now if a small startup has a great idea they will get bought out by big tech which will integrate (or kill) their tech. But now with LLMs can the likes of OpenAI just tell their model to make something that works similar to X (such as this project) and then get round copying laws and negate being behind the curve?
EDIT: switched to the correct spelling of copyright.
1. https://news.ycombinator.com/item?id=34080326
[0]https://en.wikipedia.org/wiki/Copyright
Based on my experience, sampling rate can be tricky if what you are looking for lasted less than interval period.
you are pirating first-release movies for commercial purposes?
(& by "this" I mean an approximate AI search for photos & videos - I can't account for the "every frame", nor for the comparative search quality)
otherwise, I can just make my own with my own ai. why consume someones slop when I can eat my own.