AI
ChatGPT Spots Lost Clogs in Ujjain Temple Shoe Mountain
Shubhang Borkar used ChatGPT vision to locate his clogs among over 1,000 pairs outside Ujjain’s Mahakal Temple, proving everyday AI utility beyond the laughs.
Shubhang Borkar lost his clogs in a sea of more than 1,000 pairs outside Ujjain’s Mahakal Temple and turned ChatGPT into a live scanner. The Goa-based cybersec explorer uploaded photos of the racks plus a reference shot of his footwear until the model flagged the exact spot.
His Instagram reel, posted August 12 and quickly past 35,000 likes, shows the full loop from frustration to confirmation. What looks like a joke is a clean test of vision models in uncontrolled chaos.
The clip works because the problem is ordinary and the fix is ordinary too. A phone, a consumer chatbot, and a few minutes of uploads replaced a long manual hunt through mixed piles. That combination is what made the story travel.
How the Phone Became a Shoe Scanner
Borkar first filmed the endless rows of footwear to show the problem. Distinguishing one pair from another by eye was nearly impossible in the crush.
He opened ChatGPT, uploaded a clear photo of his own clogs, and explained the scene: Mahakal Temple, lost footwear among the racks. Then he sent section after section of the piles.
The model studied the images and directed him to the right side of the rack, around the middle of the pile just below it. He checked. The clogs were there. He told the chatbot they made a good team.
- Reference image of the exact clogs first
- Context prompt naming the temple and the loss
- Sequential rack photos for progressive scan
- Specific directional guidance from the model
- On-site confirmation by the user
The whole process took minutes on a phone. No special app, no fine-tuning, just the consumer ChatGPT interface with image upload.
Order mattered. The clean reference went first so the model had a fixed target. Context named the place and the loss so the task was framed. Only then did the messy rack shots arrive, one section at a time, so comparison stayed local instead of drowning in the full pile.
Each reply narrowed the field. Spatial phrases replaced vague guesses. The user walked to the indicated zone, looked once, and closed the loop. That human check at the end kept the system honest.
Why a Temple Rack Is the Perfect Stress Test
Mahakal Temple draws massive crowds. The official site lists 1,50,000+ daily visitors. Footwear comes off before entry. Official stands and informal piles fill fast.
Hundreds to more than a thousand pairs sit in open air, mixed styles, similar colors, dirt, overlapping straps. Lighting changes. People shift shoes. Manual search burns time and patience after darshan.
Borkar, who calls himself a hacker on weekdays and explorer on weekends, treated the mess as a computer-vision challenge instead of a chore. The setting is noisy, unstructured, and high-stakes for the owner. That is exactly where many lab demos fail.
Temple scale snapshot
- 150,000+ daily visitors claimed by the temple management site
- 1,000+ pairs reported in the viral shoe piles
- 5 daily aartis and heavy peak-hour turnover
- Open racks with free or low-cost storage near entrances
A successful match here is not a staged kitchen counter. It is real visual noise.
Peak hours compound the mess. Five daily aartis pull waves of visitors through the same entrances, so racks turn over while people are still inside. A pair left in one place can drift as later arrivals shove new footwear into gaps. Memory of “somewhere on the left” fades fast under that churn.
Open storage keeps entry quick and cheap. It also leaves ownership marks thin. Similar colors and overlapping straps erase the cues a hurried owner expects to trust. The stress test is not only volume. It is volume plus motion plus sameness.
What the Model Saw
OpenAI’s vision systems let models analyze images for objects shapes colors and textures. Text in frames is readable. Multiple images can be fed in one conversation.
Borkar gave it a clean reference and messy candidates. The model compared distinctive features of his clogs against the clutter and returned spatial language a human could act on immediately.
Limitations still exist. Crowding, similar items, and poor angles can confuse systems. Here the reference photo plus iterative shots gave enough signal. The directional answer (“right side… middle… just below”) shows the model built a rough spatial map from the uploads.
This is the same capability used for document reading, product ID, or accessibility tools, applied without ceremony to temple footwear.
Feature matching does the quiet work. Shape of the clog body, strap layout, color blocks, and surface wear become anchors. The reference locks those anchors. Candidate frames are scored against them rather than described in isolation. When enough anchors align in one zone, the model can point instead of merely label.
Spatial language is the bridge to action. “Right side,” “middle of the pile,” and “just below” turn pixels into a walking instruction. Without that step, a correct visual match would still leave the owner scanning the same wall of shoes. With it, the phone shortens the last meters of search.
The Crowd Saw More Than a Gag
The video crossed 2 million views across platforms. Reactions mixed delight with recognition.
AI is solving problems nobody ever imagined it would.
GrowthSchool founder Vaibhav Sisinty wrote that after sharing the clip to nearly 100,000 views. Others asked whether they would have thought of the method at all. Jokes flew about simply grabbing the nicest pair left or trying Gemini Live instead.
One commenter recalled a husband using the same photo-and-ask trick to find a specific book on crowded sale shelves. The pattern is clear: people are already treating phone vision as a search engine for physical stuff.
That quiet shift matters more than the laughs. When a cybersec hobbyist reaches for ChatGPT the way earlier generations reached for a flashlight, the tool has left the demo stage.
Delight and practical envy arrived together. Viewers laughed at the temple chaos, then asked why they had never tried the same upload loop on their own messes. The book-shelf story shows the method was already leaking into other crowded scenes before the reel made it famous.
Jokes about taking the nicest pair left or switching to Gemini Live only underline the point. People now assume a vision model is a reasonable first tool when the eye fails. The gag lands because the habit is forming.
Everyday Chaos Meets Multimodal Models
Similar stories keep surfacing. Lost items in parking lots, specific products on warehouse shelves, landmarks in travel photos. Each one is small. Together they map the expansion of reliable visual understanding into ordinary life.
New studies on how ChatGPT reshapes thinking focus on reasoning and writing. The clog hunt shows a parallel track: perception. Users offload the visual search load and keep the decision and action.
Borkar did not wait for a dedicated “find my shoes” feature. He composed a prompt and fed images. That improvisation is the real signal. Consumer models are flexible enough that motivated users invent the applications.
| Approach | Time risk | Success factors |
|---|---|---|
| Manual rack search | High in peak crowds | Memory of location, distinctive marks |
| ChatGPT vision scan | Minutes of uploads | Clear reference photo, iterative sections, good lighting |
| Dedicated temple tagging | Depends on staff | Numbered tokens or lockers |
Temples already offer stands and sometimes tokens. The AI method fills the gap when those systems overflow or when memory fails.
Perception offload is different from asking a model to write or plan. Here the human still decides what counts as “my pair” and still walks to the rack. The model only carries the heavy compare step across hundreds of near-identical shapes. That split keeps control with the owner while cutting search time.
Improvisation spreads faster than product roadmaps. A prompt plus image upload needs no temple partnership and no new install. Wherever a phone camera can see the mess, the same loop can start. Parking lots, warehouse aisles, and market stalls are only different backdrops for the same compare-and-point pattern.
How Overflow Racks Leave a Practical Gap
Stands and tokens work until they do not. Free or low-cost open racks near entrances absorb the daily surge tied to 150,000+ visitors, yet the viral piles still held more than 1,000 pairs in open air. Overflow is the normal failure mode, not a rare glitch.
When tokens run short or informal piles grow beside official stands, memory becomes the only index. Peak-hour turnover after successive aartis scrambles that index. A visitor who trusted “I will remember the corner” meets dirt, shifted straps, and look-alike footwear instead.
- Entry rush – footwear comes off fast as crowds push toward darshan
- Rack fill – official stands and informal piles take the load without much labeling
- Peak churn – later arrivals shift pairs while owners are still inside
- Exit search – manual scanning or a phone vision loop becomes the recovery path
Dedicated tagging still wins when staff and hardware keep up. Numbered tokens and lockers remove the visual hunt entirely. The ChatGPT path matters in the gap between that ideal and the open rack reality Borkar filmed. It does not replace temple systems. It covers the overflow those systems cannot always hold.
Time risk tells the story cleanly. Manual search stretches under peak crowds. A vision scan costs minutes of uploads when the reference is clear and lighting holds. Tagging depends on staff presence. Owners now have three levers, not one, when the pile looks impossible.
From Viral Clip to Pocket Habit
Borkar’s video lands while OpenAI continues pushing image and multimodal tools, including enterprise sides such as ChatGPT Work enterprise features. The consumer version already travels in millions of pockets.
The next wave of similar clips will feel less surprising. Someone will use it for a missing bag at a station, a specific spice jar in a crowded market, or a child’s toy in a park. Each success lowers the barrier for the next person.
The Mahakal pile was never designed as an AI benchmark. It became one because a visitor with a phone decided the model was ready. The clogs came home. The larger point stayed: vision that works in the wild is no longer a lab claim. It is a weekend travel story.
Pocket habit forms through repetition, not announcements. After a temple reel, a station bag, and a market jar, the upload reflex starts to feel as plain as turning on a torch. Enterprise multimodal work may harden the same skills for offices. The consumer path is already teaching people to aim a camera at physical clutter and ask.
- Temple footwear piles under heavy daily visitor load
- Crowded sale shelves and a recalled book hunt
- Parking lots, warehouse shelves, and travel landmarks in parallel stories
- Future clips aimed at bags, spice jars, and toys in open spaces
None of those scenes was built for a benchmark. Each becomes one the moment a user treats the model as ready. Borkar’s weekday hacker habits and weekend explorer streak simply made the trial visible. The confirmation on site, the team joke with the chatbot, and the reel’s path past 35,000 likes into multi-million views turned a private fix into a shared pattern.
Frequently Asked Questions
How exactly did ChatGPT identify the specific clogs?
Borkar first uploaded a clear photo of his own pair as reference, described the temple location and the loss, then sent multiple photos of different rack sections. The model compared visual features and returned a spatial direction to the right side, middle of the pile just below.
How many pairs of footwear are typically left outside Mahakal Temple?
Reports around the incident describe more than 1,000 pairs in the piles Borkar photographed. The temple itself records over 150,000 daily visitors, so footwear volume scales with peak darshan times and festivals.
What visual details can current ChatGPT models detect in photos?
They can identify objects, shapes, colors, textures and text inside images, and they accept multiple images in one conversation for comparison. Spatial language and feature matching work when a clean reference is supplied alongside candidate scenes.
Is this method reliable for other lost items in crowds?
Success depends on a distinctive reference photo, decent lighting, and iterative shots that cover the search area. Identical mass-produced items or heavy occlusion lower the odds, yet users already report parallel wins with books on sale shelves and other household searches.
-
AI2 months agoFable 5 and Mythos 5 Return as US Lifts Anthropic Export Controls
-
AI2 months agoOracle Cuts 21,000 Jobs in a Year, Cites AI in 10-K Filing
-
AI2 months agoSpaceX’s Google Deal Turns a Rocket Company Into a Cloud Landlord
-
GAMING2 months agoCD Projekt Red Co-CEO: Redemption Arc Isn’t Done, Witcher 4 in 2027
-
CRYPTO2 months agoXPL Rallies 30% Ahead of Plasma One Card Tier Launch
-
NEWS2 months agoGoogle Search Profiles Build a Follow Graph Inside Discover
-
APPS2 months agoDGO App Brings Rs 549 Mobile Pass for FIFA World Cup 2026 in Nepal
-
AI2 months agoMoonshot AI Targets $30 Billion in China’s Fastest AI Funding Sprint
