
Similar blog to Huawei Tech4City - Synthetic Network Generation for Conversations Simulation
I've been competing in the HCMC AIChallenge again this year. This is a brief writeup of the project, and some details I liked while working on it.
Repo: YangTuanAnh/VideoWavelet
Preprocess video by downsampling frames using PySceneDetect "using rolling average of differences in HSL colorspace combined with thresholding to detect shot changes", and subtitles using Whisper for Vietnamese speech. Frames are then encoded using SigLIP2 embeddings.
VQA Module uses Gemini-3.5-Flash. TRAKE Module proposes a temporal gap regularization term for reranking. UX includes text-image search, image-image search, subtitle filtering, video reordering and filtering.
Testcase generation scripts uses Gemini-3.1-Flash-Lite to sample frames at random for query descriptions, depending on KIS, VQA, TRAKE mode.
The goal of this year's submission is to constraint myself to using as little compute possible, so that people have something to build up on. Instead of TransNetV2, the most popular option, I went for CPU-based detection. PySceneDetect took 30s per video on CPU, compared to TransNetV2 taking the same time on a T4 GPU (Google Colab). Alot more tooling was available from PySceneDetect to tweak sampling rates, thresholds and scaling options.
Also thematically, I experimented with Moonshine, specifically their Vietnamese model in Flavors of Moonshine: Tiny Specialized ASR Models for Edge Devices. Let's say it DID perform on the CPU, but the numbers were looking suspicious, given the poor output compared to the reported numbers. Had no choice but to revert back to Whisper.
Here's the prompt for the KIS frame description:
Describe this image for image retrieval, ignore the fact they come from a news segment.
Return ONLY valid JSON:
{"query": "<description>"}
And for VQA frame description:
Provide a question and answer for visual question answering.
The question should also include some description for image retrieval.
Ignore the fact they come from a news segment.
Return ONLY valid JSON:
{
"question": "<description>",
"answer": "<answer>"
}
TRAKE was a bit harder to get right to look like the typical AIC query.
These {len(images)} images are consecutive scenes from a video.
Describe the event step-by-step as a temporal retrieval query.
Focus on:
- who or what is present
- what actions occur
- how the actions progress over time
The query should describe the full sequence rather than
individual frames, ignore the fact they come from a news segment.
Example:
"a woman walks into a room, picks up a book,
sits on a couch, and starts reading"
Return ONLY valid JSON:
{{
"query": "<temporal description>"
}}
The naive algorithm would be to perform KIS queries for every step and match tuples for frames within the same video and satisfy a temporal order. This would've let through matchings where frames were too far away from each other, so a temporal gap penalty term was added. Here's a snippet from utils/server.py
gaps = [
sequence[i + 1]["scene"] - sequence[i]["scene"]
for i in range(len(sequence) - 1)
]
temporal_distance = sum(gaps)
total_score = sum(r["score"] for r in sequence)
score = total_score + 0.01 * temporal_distance
It took a while to realize how easy it is to get the video before figuring out the exact frame. Having a way to look before/ahead in the video was extremely useful to add onto the candidate pool. Image-image search is also useful in this regard since similar visuals in the video can show up at different timestamps.
There was a deliberate choice to not include popular tools such as search by regions, color fingerprinting, object tagging, since SigLIP2 already performs well enough.
Instead of having to copy back-and-forth to an external VLM to check on the frame for VQA, it would save me some time to directly ask on the site via API query.



All code and scripts are entirely human-written - except for shadcn's UI components. The use of LLMs are solely for testcase generation and VQA.
