AI for Small and Medium Businesses (SMBs) Today

The original ChatGPT-4V model has been seamlessly integrated into the current ChatGPT-4o model, eliminating the need for a separate vision-specific variant.

My first experience with the ChatGPT VLM model was while assisting my daughter in preparing for a difficult physics exam in 2024, required for university admission in the Maastricht, in the Netherlands. The exam consists of five complex, multi-page physics problems.
While chatbots like ChatGPT are excellent at predicting the next word based on context, they typically struggle with tasks that require reasoning. Initially, I tried uploading PDF files of previous exams and asking ChatGPT to solve them, but this approach didn't work. Then, almost by chance, I converted the PDF files to PNG files and uploaded those instead. To my surprise, ChatGPT correctly solved about 80% of the questions using just a short prompt. Most errors were due to copying incorrect numbers from the image, but the reasoning behind the solutions was accurate for the majority of the physics problems.
I have been extensively using ChatGPT for various vision tasks. I recommend testing the capabilities of ChatGPT by uploading one of the four images below and asking questions such as:
- What is in the fridge and provide a few suggestions for lunch today
- What is the diagnosis of the patient
- Explain what this code does
- Explain this image
I suggest trying to solve some complex math and physics problems using OpenAI's reasoning models: which are specifically designed for multi-step reasoning tasks.

What is a Vision Language Model?
There are over 15,000 large language models (LLMs) available today, including OpenAI's GPT series Meta's Llama family, Anthropic's Claude models, Google's Gemini models, and emerging reasoning-focused models that combine language understanding with advanced problem-solving capabilities.
Vision Language Models (VLMs) are multimodal AI systems that combine a large language model (LLM) with a vision encoder, enabling the LLM to see.
VLMs are versatile, context-aware, and ideal for complex tasks such as virtual sales assistants, where the virtual sales assistant can respond to both visual and verbal commands. These models are particularly valuable in marketing automation campaigns and sales automation processes, where the virtual sales assistant for example can sort items based on appearance and verbal guidance, enabling more sophisticated customer engagement strategies.
It's essential to distinguish Vision Language Models (VLMs) from image generation tools like DALL-E, Midjourney, and Stable Diffusion. A VLM is trained on a large dataset of image-caption pairs and does not generate images from a text prompt. A VLM is a (prompt + image) to text model, and outputs text only.

VLM Use Cases
VLMs are quickly becoming the go-to tool for all types of vision-related tasks due to their flexibility and natural language understanding. VLMs are capable of tasks such as image analysis, visual Q&A, image summarization, solving complex math and physics problems. A VLM can also assist the visually impaired by analysing images and generating descriptive text, which can then be read aloud using text-to-speech technology: image + prompt to text to audio.

With vast amounts of video being produced every day, it's infeasible to review and extract insights from this volume of video that is produced by all industries. VLMs can be integrated into a larger system to build visual AI Agents capable of detecting specific events when prompted.

These systems could be prompted to detect safety breaches, detect empty shelves, detect traffic incidents, etc.
This visual AI Agent can handle diverse inputs and automate tasks that traditionally required manual monitoring. It detects events, identifies activities, and provides actionable insights or outputs based on the analysed visual data, demonstrating its broad application potential in surveillance, sports analysis and emergency response.
Advances in deep learning, reinforcement learning, and multimodal AI have greatly improved the perception, reasoning, and decision-making capabilities of virtual sales assistants. VLMs allow these assistants to understand both visual inputs and natural language instructions, enabling rich contextual understanding. By analyzing images and responding to prompts, VLM-powered assistants can recognize visual elements, interpret sales-related content, and carry out complex tasks with precision.
For example, a VLM-powered sales assistant can look at a screenshot of a CRM dashboard, spot high-value leads based on visual cues (like deal size or activity level), and follow an instruction such as "send a follow-up email to the top 3 leads."
Similarly, Salesforce Einstein could analyse a visual sales pipeline chart and respond to a prompt like "highlight stalled deals from last quarter." It combines visual analysis with natural language understanding to quickly surface key insights and take action, just like a smart assistant would.
VLM Challenges
VLMs are maturing quickly, but they still have limitations, particularly around spatial understanding and long-context video understanding. Long video understanding is a challenge due to the need to take into account visual information across potential hours of video to properly analyse or answer questions. Like LLMs, VLMs have limited context length meaning, only a certain number of frames from a video can be included to answer questions.
Implementing VLMs in virtual sales assistants presents several challenges, including high costs, technical complexity, and ethical concerns. Integrating vision and language capabilities requires sophisticated models and complex engineering. Training VLMs involves large image/caption datasets and high computational power. Organizations implementing these systems often benefit from AI training programs and specialized AI consulting services to navigate the technical complexity and maximize ROI.
The increased automation of tasks traditionally performed by humans raises ethical issues around job displacement and labour impacts. It should be no surprise that some countries are evaluating the 4-day workweek and the universal allowance for people that fall through the cracks.
As VLM-enabled virtual sales assistants perform more complex tasks, there's a risk of replacing jobs in marketing and sales departments including automated content generation, marketing campaign delivery, sales outreach, and service industries, leading to economic and social concerns.
Organizations implementing VLMs must also consider data privacy and copyright implications, and especially when processing visual content that may contain sensitive information.
VLMs are incredibly powerful and often built into large language models. If you're not using them at least 50% of your AI work, you're missing out. There is no need for special tools, just use them directly in platforms like ChatGPT, Claude, or any modern LLM. To get proficient in image to text applications, it typically takes a short, one time training session of 90 minutes.
Rainmakers SG helps small and medium businesses design safe and scalable Agentic AI systems that provide immediate ROI!



