Model Matters

Getting the best results when exploring AI

Artificial Intelligence is a revolutionary breakthrough. It can help researchers understand qualitative data in new, interesting, powerful ways. Occasionally, as I develop and refine new features for Transana, I discover subtleties to working with AI, which I share here.

Transana allows researchers to explore individual snapshots with AI. It also facilitates AI exploration of text in Collections. Today, I am working on adding AI processing of Snapshots in Collections, which should allow exploring more than one still image at a time. There’s a mechanism in Ollama that should be able to handle multiple still images. So I created a Collection with the two still images below and I modified Transana’s source code accordingly.

Assyrian Lemmasu
Assyrian Lion

But the AI exploration did not work in my preliminary tests. The results of my prompt didn’t make any sense using the “gemma4:26b” model. (Gemma4 is a newer model that’s done really well with exploring text, and it’s become my “default” model at the moment.)

The first thing I did was modify my prompt so that the AI would describe each image. This showed that the AI didn’t appear to “see” the second picture.

The first image depicts a large, ancient stone relief of a lamassu, an Assyrian protective deity featuring a human head, a bull or lion’s body, and large feathered wings. The sculpture is highly detailed, showing a finely textured beard, a distinctive headdress, and intricate patterns on the wings and musculature. Cuneiform inscriptions are visible carved into its side. The artifact is displayed in a museum setting against a white wall, with another smaller relief carving partially visible to its left.

The second image does not contain any recognizable objects or subjects; instead, it consists entirely of dense, horizontal parallel lines of varying widths and colors. These lines create a pattern of digital noise or “scanlines,” transitioning through shades of dark brown, tan, black, and white. It appears to be an abstract texture or a result of digital corruption/static.

So I started experimenting. I checked that the images were loading correctly and were available to Ollama. Each image worked separately; I only had problems when I tried to see both images. I increased the size of the context window. Then I tried a different model, llava:7b. It’s an older model, but has a good reputation online for processing images. (Note: llava:7b has not been a top performer historically in my single-image tests in Transana.)

The descriptions of the images produced by llava:7b, as shown below, are not very good, but it is clear that this model does indeed “see” both images. That suggests my code is working but that gemma4:26b is not able to process multiple images. Additional tabs below show the responses from other AI models to the same prompt and same images.  Click the model name to see that model’s response.

The first image displays an ancient-looking sculpture of a creature that appears to be a combination of human and animal features—possibly a mythical or fantastical hybrid being, often referred to as a griffin or chimaera. This creature has the body of a humanoid figure with the head of an eagle. The sculpture is positioned in front of some form of display or exhibit, likely within a museum setting, as suggested by the glass case and the wall with information boards.

The second image features another ancient-style statue, but this one appears to be of a different creature entirely—a lion with a humanoid head. The sculpture is larger than the one in the first image and also resides within a museum exhibit, indicated by the display case and informational signage.

The first image shows an ancient stone sculpture of a lion, which appears to be from the Babylonian period. It has detailed carvings on its body, including lines that might represent muscles or fur. The head of the lion is stylized with intricate details, and it seems to be in a state of roar or shouting. There are other smaller pieces of sculpture visible in the background, but they are not detailed enough to describe accurately.

The second image depicts a different stone lion statue, which appears to be from an earlier period, possibly Sumerian or another ancient civilization’s artwork. This lion has more abstract and simplistic features compared to the Babylonian sculpture. It does not show as much detail in its facial expression or body muscles but is equally evocative of a powerful figure.

**Image Description 1:**

This image showcases a monumental stone sculpture, believed to be one of the Lamassu from the ancient city of Nineveh in Assyria. The statue is depicted as a hybrid creature – it has the body of a bull or lion, wings resembling an eagle, and a human head with a stylized headdress. The figure is highly detailed, emphasizing muscular definition and intricate feathering. It stands upon a rectangular block covered with cuneiform inscriptions, suggesting its original placement within a grand palace or temple complex. The lighting highlights the texture of the stone, revealing the sculptor’s meticulous work, and casts shadows that accentuate its imposing size and powerful presence.

**Image Description 2:**

This image presents another nearly identical Lamassu sculpture also originating from Nineveh. It is presented in a museum setting surrounded by glass display cases to protect it. The statue shares the same hybrid characteristics as the first: a leonine or bull body, feathered wings, and a human head adorned with a headdress. This one appears even larger than the previous image due to the perspective and surrounding display. Like the earlier photograph, detailed modeling of musculature and feathers is clearly visible. It’s positioned on a similar rectangular base inscribed with cuneiform text.

**Image 1 Description:**

The first image showcases a colossal sculpture within a museum setting. The centerpiece is a Lamassu – a mythical creature from Assyrian mythology. It’s depicted as having the body of a bull or lion, the wings of an eagle, and the head of a human wearing a horned headdress, characteristic of royalty. The sculpture appears to be carved from stone (likely gypsum alabaster), displaying intricate details in its musculature, feathered wings, and facial features. Behind the Lamassu is a relief panel with further carvings, indicating this likely served as part of a larger architectural decoration. A museum label (“68”) is visible near the top edge, signifying it’s an exhibit piece. The lighting highlights the texture and grandeur of the sculpture.

**Image 2 Description:**

This image features another large-scale sculptural work in what seems to be the same or a similar museum environment. It portrays a massive lion statue with its mouth wide open in a roaring expression. Like the previous one, it’s made from stone (again likely gypsum alabaster) and possesses impressive detail. The lion’s fur is rendered realistically, conveying strength and power. There are inscriptions carved into the base of the sculpture, suggesting they might be cuneiform text, an ancient writing system used in Mesopotamia. It sits on a modern display platform enclosed by glass to protect it from visitors and for viewing purposes. A label (“68”) can also be seen in this image.

I apologize, but the two images you wish me to describe and compare were not included in your request.

Only one image was provided with your request. It is a photograph of a large stone sculpture in a museum setting.The image features a massive **lamassu**, a protective deity from ancient Mesopotamia (specifically Assyrian). The sculpture depicts a creature with the head of a human, the wings of a bird, and the body of a bull or lion. The piece is meticulously carved from stone, showing intricate details such as a stylized beard on the face, individual feathers on its large wings, and muscular proportions in the legs. In the background to the left, another smaller relief sculpture is visible, and several informational plaques can be seen on the walls. The presence of glass barriers and museum signage indicates it is an exhibit in a gallery like the British Museum.

**First Image**: This photograph features an ancient Assyrian *lamassu*—a composite mythical creature with a human head, bull’s body, and large wings. Carved from stone (likely limestone or alabaster), it showcases intricate details: the human face wears a distinctive headdress and has elaborate facial hair; the wings are rendered with feather-like precision; and the bull’s body exhibits muscular contours. The sculpture is displayed in a museum setting, accompanied by smaller relief panels depicting similar mythological figures, with soft lighting emphasizing its weathered stone texture and historical significance.

**Second Image**: This image presents another ancient stone sculpture, specifically a winged lion (or a stylized lion variant). Mounted on a pedestal within a museum environment, it features detailed fur patterns and an open mouth, suggesting aggression or power. The surface shows signs of erosion from age, typical of ancient artifacts, while the background includes architectural elements like columns and windows, situating it in a curated exhibition space.

The first image depicts a large, ancient stone sculpture of a mythical creature with the body of a lion and the head of a human being. The figure appears to be wearing a helmet and has detailed facial features, suggesting it may represent a deity or king from an ancient civilization. The craftsmanship is intricate, showing the attention to detail in the carving of the creature’s mane and the textured surface.

The second image shows another stone sculpture, this time of a lion standing on its hind legs with its front paws resting on what appears to be a base or platform. The lion has a fierce expression, with its mouth open as if roaring. The texture of the stone is rougher compared to the first image, and the carving style seems more primitive.

1. **First Image:**

The first image depicts an ancient stone relief sculpture of a mythical creature known as a lamassu. This figure has the body of a bull or a lion, wings of an eagle, and a human head wearing a horned helmet. Lamassus were prominent in Mesopotamian art and were believed to be protective deities.

2. **Second Image:**

The second image shows another ancient stone sculpture, but this one is of a large lion. This sculpture appears to be from the same Mesopotamian or Assyrian period as the first image. The lion is depicted roaring, with its mouth open and teeth visible, showcasing detailed craftsmanship.

### **Image 1 (Top-left)**

This image depicts an ancient Mesopotamian statue featuring a human-headed winged lion or bull, likely representing a **”lamassu”** or protective deity. The figure has the body of a bull and the head of a bearded man wearing a horned headdress, symbolizing divinity in Assyrian art. Intricate cuneiform inscriptions are visible on its body, adding to its grandeur and sacred significance.

### **Image 2 (Top-right)**

This image shows another ancient Mesopotamian sculpture—a massive stone lion with an open mouth, likely part of a palace or temple guard. The craftsmanship is detailed, emphasizing the lion’s muscular build and fierce expression. This type of statue was often placed at entrances to protect sacred spaces.

1. **First Image:**

This image shows part of a large stone relief featuring a mythical hybrid figure known as a lamassu (or shedu in some contexts). The lamassu is a protective deity from Mesopotamian mythology, often depicted with the body of a lion or bull, human head, and sometimes wings. This particular sculpture has a human-like face, a headdress, and the muscular, powerful body of a winged bull.

2. **Second Image:**

This image shows another stone relief of a lamassu, but this time it is specifically in the form of a lion. The carving captures the majestic posture of a lion with detailed musculature and a fierce expression, common in ancient Assyrian art used to symbolize strength and protection.

We get very different responses with different models to the same prompt and the same images. In this particular example, only some models give us accurate descriptions of both images. It is interesting and surprising to note that the newer gemma4:12b model was not able to process two images where the older gemma3:12b model was. It is also interesting, although not surprising, to not the differences between gemma3:4b, where the image descriptions are inaccurate, and gemma3:12b, where the image descriptions are much more detailed and accurate.

I draw three important conclusions from this experience:

  • When working with images, it helps to have the AI describe each image to ensure it is “seeing” what it is supposed to. The AI model will not always tell you when it is not set up for processing images following the Ollama rules for multiple images.
  • It can be helpful to submit the same AI exploration prompts and data to several different models. You can expect different results from different models.
  • AI results are suggestions, not reliable conclusions. They are not always accurate, and it can be hard to tell at a glance when they are not. It is vital to check all AI results against your data to ensure that they accurately reflect your data.