Scylla
Get a Quote
Building an Identity-Aware Multi-Object Tracking System

Building an Identity-Aware Multi-Object Tracking System

Detecting a person in a video frame is only the beginning. To follow that person through a scene, a system must keep connecting new observations to the same individual - even when people overlap, change direction, or become partially hidden. That continuity is central to our research at Scylla: building a tracking system that combines visual identity with an understanding of movement over time.

Our system has achieved first place in the official MOT17 and MOT20 competitions hosted on CodaBench. These milestones follow Scylla’s first-place result in the COCO Detection Challenge (Bounding Box) in 2025, extending our benchmark achievements across object detection and multi-object tracking.

Behind the rankings is a broader engineering effort: improving how a system represents a person’s appearance, identifies plausible matches, and uses previous observations to preserve a consistent track.

Learning to recognize the same person

Person re-identification, or ReID, helps a system determine whether different observations show the same individual. In this context, identity means maintaining the same person’s identity within the tracking system; it does not require knowing their name.

This is challenging because a person’s visual appearance can vary considerably. A useful representation must also remain effective when the visual environment differs from the data used to train the model. This ability to work across different visual domains is known as cross-domain generalization.

Group army men saluting

To address these challenges, we developed a multi-branch ReID architecture. The network learns several complementary visual characteristics of a person through separate processing branches. Together, these branches provide information that supports identity matching.

We also developed a new loss function for joint ReID and attribute learning. A loss function guides training by defining which errors the network should reduce. Our approach brings identity and attribute information together during training, helping the network learn how these complementary signals relate to the same person.

Supporting this work is an automatic data annotation pipeline that generates attribute labels at scale, reducing dependence on manual annotation. The resulting representation showed strong cross-domain performance in our Market-1501 evaluation. The research objective is to develop identity representations that remain useful as the visual domain changes.

Choosing plausible matches before comparing appearance

A second challenge is association: deciding which newly detected person belongs to which existing track.

Many tracking systems rely heavily on intersection over union, or IoU, which measures how much two bounding boxes overlap. A bounding box is the rectangle a detector draws around an object. Overlap provides a useful spatial signal, but it cannot by itself establish that two boxes belong to the same person.

Group army men saluting

Consider two people passing close to each other. Their boxes may overlap, and one person may briefly obscure the other. Selecting a match primarily because its box overlaps most can connect a track to the wrong individual.

Our approach first considers which detections are plausible candidates. It uses the object’s velocity, direction of movement, previous trajectory, bounding-box position, and aspect ratio - the relationship between the box’s width and height. It also considers how that shape changes over time.

These changes provide useful context. As a person approaches the camera, turns, or becomes partially obscured, the geometry of the visible box can change. Interpreting that geometry alongside motion and trajectory helps the system narrow its search for a matching identity.

The association process therefore asks which detections are consistent with the object’s motion and history. ReID then helps determine the corresponding identity among those candidates, drawing on historical appearance information. Spatial plausibility and visual identity work together in the matching process.

Connecting the research across the pipeline

This work spans both model development and tracking. Automatic annotation supports the training data; the multi-branch architecture and joint loss shape the learned representation. During tracking, object detections are assessed using motion, trajectory, and geometry before historical ReID information helps resolve identity.

Each part addresses a different source of uncertainty. Detection establishes where an object appears. Motion and geometry help identify plausible continuations of its track. ReID contributes evidence that the observation belongs to the same person.

Group army men saluting

In the MOT17 announcement, IDF1 score of 87.35 and 708 identity switches for SCYLLA-MOT are reported. IDF1 reflects identity consistency during tracking; an identity switch occurs when the tracker incorrectly changes the identity associated with an object. These results provide a measurable way to assess the continuity our approach is designed to improve.

Why these benchmarks matter to security users

COCO, MOT17, and MOT20 address complementary capabilities that security video analytics depend on. COCO evaluates how accurately objects are detected and located in images, while the MOT competitions evaluate tracking across video. MOT17 includes challenges such as crowded scenes, occlusion, and changing camera perspectives.

The practical implication is that accurate detection provides the starting observations, while consistent tracking helps operators follow the same person through an unfolding event. Together, these capabilities can support a clearer account of movement and reduce confusion caused by fragmented or incorrectly linked tracks. Benchmark results provide evidence of these underlying capabilities under defined evaluation conditions; they do not measure every aspect of performance at an individual security site.

Our continuing research focuses on the relationship between appearance, movement, and history. By developing the algorithms behind each stage, we aim to make identity continuity more reliable when the scene becomes difficult - the point at which security users most need a coherent view of what is happening.

About the Author

Zhora Gevorgyan

Zhora Gevorgyan

Lead Computer Vision Engineer, Scylla AI

Zhora Gevorgyan is a computer vision engineer, researcher, and AI architect widely recognized for innovation in real-time object detection and machine learning. As the Lead Computer Vision Engineer at Scylla, Zhora Gevorgyan has been behind the core technical breakthroughs that power Scylla's platform, specializing in next-generation threat detection and public safety infrastructure. He is best known as the creator of ScyllaNet which ranked 1st in 9 out of 12 evaluation metrics on the prestigious COCO object detection leaderboard. He is also the author of the revolutionary SIoU (Scylla-IoU) loss function. Together, these innovations have established new global benchmarks for architectural efficiency and algorithmic precision.

Learn More

Stay up to date with all of new stories

Scylla Technologies Inc needs the contact information you provide to us to contact you about our products and services. You may unsubscribe from these communications at any time. For information on how to unsubscribe, as well as our privacy practices and commitment to protecting your privacy, please review our Privacy Policy.

Related materials

Scylla is AICPA certified
Scylla is ISO certified
Scylla is ASPP certified
GDPR compliant

Copyright© 2026 - SCYLLA TECHNOLOGIES INC. | All rights reserved