Local Qwen Does Well With Images
EricMesa
- 3 minutes read - 437 wordsAfter spending 2 posts talking about how badly these local AI models are doing on my machine (here and here ), I did want to show a example of the model performing surprisingly well. Using qwen 3.5 I asked it to describe the following image:
A photo taken a long, long time ago
Here is the description that qwen wrote after about a few minutes:
This is a vertical black-and-white photograph featuring a group of six young Vietnamese women posing on concrete steps outside a building. Here are specific details about the scene:
The People and Attire:
- All six women are dressed in traditional áo dài, which consists of long, tunic-style tops worn over matching trousers. The fabric appears to be white or light cream.
- They have dark hair styled straight and worn down past their shoulders.
- Most of the women are smiling at the camera.
Accessories:
- Several women are holding traditional conical hats (known as nón lá). One woman on the far left holds one under her arm, while another in the center foreground holds a large hat out with both hands. The two women furthest to the right appear to be carrying larger bags or perhaps folded mats/hats against their legs.
- The second woman from the left is holding a dark, rectangular handbag or briefcase against her hip.
The Setting:
- They are standing on wide concrete steps leading up to an entrance.
- Behind them is a multi-story building with architectural features that look somewhat colonial or institutional (possibly a school, government office, or apartment complex). It has columns and balconies with railings running along the upper levels.
- To the right of the group, there is a leafy tree. In the far background on the right edge, another small figure—likely a child—is visible standing near some crates or boxes in what looks like an alleyway or courtyard area.
The photo captures a candid yet posed moment of daily life, likely taken in Vietnam (possibly Hanoi) during the mid-to-late 20th century given the fashion and architecture style.
What I don’t understand is how qwen could be SO BAD at finishing its thoughts on a book summary and could be SO GOOD at describing this photo. If anyone knows, I’d love to understand why it works this way. I will say, and I knew this was the case with these LLMs - the output is not deterministic. This is similar, but not the same as when I asked it the same question yestrday. Yesterday it guessed it was from the 60s or 70s and was in Ho Chi Minh City. Today it says mid-to-late 20th century and Hanoi.