How I Used AI to Catalog 100+ Magazines (and Where It Failed)
A while back, I was trying to find a specific article in a physical magazine. I had over 100 of them, and Google was no help at all. So I figured I'd try something different: point AI at the covers and see if it could help me find what I was looking for.
That one experiment turned into a bigger question: can AI handle the boring work of cataloging books, magazines, and other physical records using OCR and text analysis?
Here's what I tried, what worked, and where I still had to keep an eye on it.
The goal
Simple: digitize a stack of magazines by pulling titles, dates, and key topics straight from photos.
Instead of typing everything by hand, I wanted AI to look at the images, read the text, and give me something structured back.
Setup
- 100+ magazines
- non latin language - Cyrillic alphabet
- light - daylight
- different magazines - red, yellow, green
- different fonts and colors - red, yellow, green
Step 1: Taking the photos
AI is smart, but it's not magic. If you give it blurry, dark, or glare-heavy photos, you're going to get bad results.
Input quality made a big difference.
- Lighting: I tried a few setups to cut down glare on glossy covers.
- Orientation: I kept titles horizontal and made sure the text I cared about was visible.
- Image quality: I played with camera and image settings. Cleaner images almost always meant better OCR.
- Orientation: There wasn't big difference for horizontal vs vertical image.
Lesson: better input means less cleanup later and higher accuracy.
Step 2: The prompt
Once I had photos, I needed a way to turn them into structured info.
I wrote a prompt asking the AI to do OCR and put the results into a Markdown table.
Here's the prompt I used:
OCR the text and return it as a Markdown table with these columns:
- Date: септември 2024
- Title:
- Topics:
- ЗНАМЕНА НА СВЕТИНИ.
- КРАЛСКАТА ГВАРДИЯ.
- ОПАЗВАНЕ НА НАСЛЕДСТВОТО.
- КОСМОС.
- МИГРАЦИЯ
Pretty simple on purpose. I just wanted the visible text pulled out and dropped into a predictable format I could copy, search, or process later.
Step 3: Trying different models
I tested the workflow with a few free AI tools:
- Claude
- ChatGPT
- DeepSeek
- Kimi
The results were surprisingly similar.
None of them were perfect all the time, and they all ran into the same kinds of problems. For this task, at least, the limitations didn't seem tied to one model.
Final results were taken from ChatGPT.
Step 4: Checking the results
The output was impressive - but definitely not flawless.
Across the images I tested, field-level accuracy (titles, dates, key topics) was around 80-90%.
That's good enough to save a ton of time compared to typing everything manually. It's not good enough to trust blindly.
What worked
The AI could usually pick out:
- Magazine titles
- Dates
- Headlines
- Topics
- A lot of the text and some layout details
Instead of spending hours transcribing covers, I could get a first draft of the catalog in seconds. That alone made the experiment worth it.
What didn't: silent failures
The worst part wasn't when the AI said it couldn't read something.
The worst part was when it read something wrong and sounded completely confident about it.
For example:
- Moon became Noon
- mission became version
Those mistakes are easy to miss because they look totally plausible.
That's one of the big differences between traditional OCR and AI-based image analysis. The AI may try to make sense of unclear text instead of just saying it can't read it.
The honest failures
To be fair, the models didn't always guess.
Sometimes they correctly said the text wasn't readable, returning things like:
not visible???
Those responses are much safer than a confidently wrong word.
For cataloging and archival work, an explicit unknown is usually better than a plausible mistake.
Input


Output
| Дата | Основна тема | Останал OCR текст |
| --- | --- | --- |
| юни 2012 | ХРАНЕНЕ | Тайна медицина · Аквариус · Сеул · Мексико · Перелман |
| октомври 2010 | ЛЮБОВНАТА ДВОЙКА | Девадаси · Райски птици · Тайван · Етнология · Стрес · Патагония · Крионика · Биоразнообразие |
| май 2015 | ГЪРЦИЯ | Светлина в Мянмар · Демeнция · Дигитална революция в Африка · Проект ICARUS · Шрауси · Химическа война · Телескопът Хъбъл |
| март 2015 | ШВЕЦИЯ | Черният дупник · Завръщането на раса · Криза в Ориента 1839 · Средновековни кули · Отровна пещера · Природна фотография |
| февруари 2015 | ДРЕВЕН ЕГИПЕТ | Мозък · Акция на Витоша · Белиз · Пустинна лисица · Дивия Запад |
| януари 2015 | ШАНХАЙ/ХОНГ КОНГ | Пещерата Шове · Мексиканският краал на дрогата · Планините на Антарктида · Дървета · Мали |
| юли 2014 | КУБА | Слонове · Абруцо · Биохакинг · Острови на подправките · Гледки отдалеко |
| октомври 2013 | РИМ | Фрайдинг · Цветове · Шамани · Северна Корея · Етиопия |
| юни 2013 | ЕВРОПЕЙСКИ ПЛАЖОВЕ | Вода · Пясък · Естетична хирургия за мъже · Ленивци · Хитлер · Етиопия |
| май 2013 | МЛАДАТА ЕВРОПА | Речна фауна в Бразилия · Пустинята Атакама · Фитнес парк в Киев · Хамбург · Алергии · нови стратегии |
| април 2013 | ЧУДЕСА | Леопарди · Курьозни фотографии на животни · Ладак · Фалшиви медикаменти · Ботаниците на Сталин · Експедиция във Венесуела |
| февруари 2013 | КИТАЙ | Земни съни · Надвитата рулетка · Съвременни еремити · Бунтът на „Баунти“ · Всевиждащи сателити |
| октомври 2012 | КУЧЕТА | Дивата Парагвая · Планински спасители · Кубинската криза · Незабравимата Сицилия · Биопиратство |
| септември 2012 | КЗПЕРА? | Спасителите на каракачанците · С колело през Китай · Царството на гъбите · Маймуните · Индианците и цивилизацията |
| март 2012 | НЮ ЙОРК | Тайгани · Полет · Алцхаймер · Австралия · Пергам · Лов на китове · Път в Непал |
| декември 2011 | МАРИЯ / БОГИНИ | Виетнам · Сурикати · Банкок · Черни мури · Световно наследство · Трансплантации · Търговия с цветя |
| октомври 2011 | ЖИВОТ И ЩАСТИЕ | Фотосинтеза · Малта · Малдиви · Ягуари · Метеорология · Ландшафт · Експедиция в Австралия |
| септември 2011 | ТАКЛАМАКАН | Фестивалът „Холи“ · Мали · Близкоизточна съдба · Версай · Мохамед · Биоразнообразие |
| юли 2011 | ТЪРГОВИЯ С ДИВИ ЖИВОТНИ | Марс · Коста Рика · Талин · Сасан-Гир · Карл V · Папуа-Нова Гвинея · Микроби |
| май 2011 | КИТАЙ | Бугараш · Медузи · Препаратори · Охлюви · Деца и медии · Тимбукту · Северна Ирландия |
| април 2011 | ПЪРВИТЕ ХОРА | Космос · Мон Сен Мишел · Леошоди · Традиционни знания · Светилища · Земноводни · Бащи · Боливия |
| март 2011 | БОЛКАТА | Кайлас · Размножаване на растенията · Корпус на мира · Джибути · Орхидеи · Златни перли · Молдова |
| февруари 2011 | 1941 ГОДИНА | Големият бариерен риф · Пакистан · Числото Пи · Кантабрия · Grand Paris · Чарлз Брюър-Кариас · Японски градини |
| септември 2010 | МОРСКИ ИЗСЛЕДВАНИЯ | Мимикрия · Балет XXL · Гърция · Ендемити · Уинице · Белинташ · Тигри · Бербери |
| август 2010 | КАТНА | Роботи · Храмът на Зороастър · Водорасли · Тунис · Бавноходки · САЩ Трой · Дървена мафия |
_Бележка: няколко думи по ръбовете/в замъглените участъци са неясни от снимката, затова са отбелязани приблизително._
What I learned
This changed how I think about using AI for this kind of work.
AI doesn't have to be 100% accurate to be useful.
If it can turn an hours-long task into a minutes-long one - and leave a human with just the errors to check - that's still a huge win.
The key is building the workflow around the tech's limits.
In this case, the workflow looked like this:
Take photos, run OCR with AI, structure the data, validate the results, then add it to the catalog.
AI does the repetitive part. A human does the checking.
Compare to old OCR
I can complete the task by using classical Python OCR or other libraries. You can find example in this article: python extract text from image or pdf
By using the AI I got:
- less control
- less efforts
By the end it saved few hours from this task.
Conclusion
AI is great for automating boring, repetitive tasks. For cataloging magazines, books, and other physical records, it can give you a massive head start.
But it's not a "set it and forget it" solution.
The biggest risk isn't an obvious failure. It's a silent failure - when the AI gives you something that looks right but isn't.
So I wouldn't use AI-generated OCR as the final source of truth for an archive. I'd use it as a fast first pass, then have a human validate it.
In other words:
Let AI do the boring work. Just don't let it have the final word.