AI-Powered Medical Video Search with Azure and NVIDIA

90%
improved search accuracy
~68%
reduced video review time
AI-Powered Medical Video Search with Azure and NVIDIA | Crunch-IS Case Study
Created an AI-powered video analytics solution on Azure OpenAI and NVIDIA H100 GPUs, delivering secure, compliant, and instant insights from medical videos for our client.
Industry:

Healthcare (Diagnostics)

Location:

UK

Duration:

6 months

Technologies
Azure Machine Learning
Azure OpenAI services
Azure Cosmo DB
Open source VLM
ASR
Neo4j
01

About the Client

The client is a healthcare service provider known for its innovation in diagnostics and physician training. With operations across multiple hospitals and research centers, the client sought an video AI-powered solution to streamline access to video-based knowledge without compromising on accuracy, performance, or data privacy.

AI-Powered Medical Video Search with Azure and NVIDIA | Crunch-IS Case Study
02

Challenge

The client’s clinical teams were burdened with reviewing hours of unstructured video recordings to locate key decisions, symptoms, or procedural steps.

Video data recorded from endoscopies, operating rooms, and training labs was stored in silos with no way to semantically query or summarize content. Manual transcription and indexing were inefficient and impractical at scale.

Additionally, all data had to be processed within secure cloud infrastructure to ensure compliance with GDPR and UK NHS data governance policies.

03

Project Scope

We developed a high-performance video AI agent solution leveraging the full-stack Azure AI ecosystem including Azure Machine Learning, OpenAI services, Cosmos DB, Blob Storage, Azure Functions, and critically, the NVIDIA H100 v5 series virtual machines. These cutting-edge GPUs provided the compute power necessary to handle real-time video analysis and inference at scale.

The goal was to enable instant search, contextualization, and summarization of medical video content; all to help clinicians save time, reduce errors, and improve training and decision-making.

Crunch-IS Case Study | Image

Video Data Structuring and Integration

Video is inherently complex, multimodal data. Our first task was to ingest and structure large archives of clinical video content (consultations, endoscopies, surgical simulations).

Our team integrated disparate data sources, automated metadata enrichment, and aligned the video corpus with the client’s knowledge systems for consistent downstream use.

Multimodal AI and Feature Engineering

We prepared detailed analysis reports covering the video corpus, data distributions, quality benchmarks, and medical context correlations. Combining clinical expertise with advanced feature engineering techniques, we transformed raw video into multimodal datasets (captions, audio transcripts, embeddings) suitable for AI/ML processing.

AI-Powered Architecture

The solution was architected around key Azure services to deliver real-time performance and scalability:

  • Azure Machine Learning – simplified access and deployment of VLMs and LLMs (e.g. Llama, Qwen, Parakeet) on NCads H100 v5 virtual machines.

  • Azure Functions – orchestrated Blob Storage events, initiated VLM and ASR tasks, managed task completion, and triggered targeted search and summarization workflows from generated multimodal captions.

  • Azure OpenAI Services + Cosmos DB (optionally Neo4j) – built a natural language video search interface, enabling clinicians to query recordings with expressions such as:

    “Show me when the surgeon identified bleeding”, “Summary of post-op instructions”.

Secure Data Processing

Since data governance is critical in healthcare, all video data is processed on-premise, either within secure GPU-accelerated Azure cloud infrastructure. This ensures compliance with GDPR and UK NHS policies while maintaining performance, scalability, and reliability.

04

Results and Achievements

By the 3d month, the video AI agent solution was fully operational. Doctors, clinicians, and researchers gained the ability to:

· Search thousands of hours of clinical video content using natural language queries;

· Instantly retrieve critical procedural highlights and patient interactions;

· Access summarized views of lengthy surgical recordings in under a minute;

· Integrate annotated transcripts and video metadata into customer knowledge based system.

The solution dramatically reduced review times, improved diagnostic workflows, and freed up valuable clinical and research staff hours. Additionally, by using custom fine tuned open source vision language models tailored to the healthcare domain, the system achieved faster and more accurate results than generalized cloud AI services.

Have a Question? Let’s Get in Touch!

Tell us what you’re building or where you’re stuck. We work with engineering and product teams on custom software, AI & ML, cloud infrastructure, DevOps, and UI/UX — from early scoping to long-term delivery. One conversation is usually enough to know whether we’re the right fit.

Email: [email protected]

    I have read and accepted the Terms of Use and Privacy Policy *

    We and our partners use technology such as cookies on our site to personalize content and ads, provide social media features, and analyze our traffic. Click “Accept” to consent to the use of this technology across the web.

    Decline