| dc.contributor.advisor | Volpe, Gualtiero <1974> | |
| dc.contributor.advisor | Ceccaldi, Eleonora <1990> | |
| dc.contributor.advisor | Varni, Giovanna <1979> | |
| dc.contributor.author | Serajoddin Mirghaed, Amirarsalan <1998> | |
| dc.date.accessioned | 2026-07-23T14:32:21Z | |
| dc.date.available | 2026-07-23T14:32:21Z | |
| dc.date.issued | 2026-07-16 | |
| dc.identifier.uri | https://unire.unige.it/handle/123456789/16495 | |
| dc.description.abstract | Abstract
This thesis proposes a computational approach to Generic Event Boundary Detection
(GEBD). GEBD involves identifying when meaningful events occur in an untrimmed
video, without predefined action classes. The proposed approach is grounded in the Event
Segmentation Theory (EST), which posits that humans detect boundaries by predicting
errors when their internal model of the ongoing situation fails.
EST is operationalised through a 15-channel feature representation aligned to EST
perceptual categories of time, space, objects, characters, and interactions. Such a fea-
ture representation is fed to three different machine-learning architectures: ReconViT-PD
(RVpd), a Vision Transformer (ViT) with Nyström-based attention, a frozen-encoder
GRU predictor (RV-PL), and a combination of the former so that the GRU receives the
sequence of the embeddings the ViT produced. Event boundaries are detected by applying
peak detection to a change-score time series. Results are compared with those obtained
by an established baseline model (BL).
The approaches were tested on two datasets with distinct application scenarios: As-
sembly 101, a multi-view video-recording dataset featuring 53 participants assembling and
disassembling 101 toy vehicles, and Breakfast, a dataset of multi-view video recordings
of cooking activities. MCC with a ±15-frame tolerance was adopted as the primary per-
formance metric. On Assembly 101, results show MCCs of 0.405 (RVpd), 0.390 (EST),
0.228 (RV-PL, tying the baseline) and 0.228 (BL). For Breakfast, the results show MCCs
of 0.593 (RVpd), 0.514 (EST), 0.425 (RV-PL), and 0.256 (BL). A cross-corpus experi-
ment was also conducted by training the models on Assembly 101 and testing them on
Breakfast, yielding MCCs of 0.577 (RVpd), 0.522 (EST grid), and 0.425 (RV-PL).
A deeper investigation of the results reveals that the per-frame change score the mod-
els emit does carry boundary information (ground-truth frames score on average 1.75
times higher than | it_IT |
| dc.description.abstract | Abstract
This thesis proposes a computational approach to Generic Event Boundary Detection
(GEBD). GEBD involves identifying when meaningful events occur in an untrimmed
video, without predefined action classes. The proposed approach is grounded in the Event
Segmentation Theory (EST), which posits that humans detect boundaries by predicting
errors when their internal model of the ongoing situation fails.
EST is operationalised through a 15-channel feature representation aligned to EST
perceptual categories of time, space, objects, characters, and interactions. Such a fea-
ture representation is fed to three different machine-learning architectures: ReconViT-PD
(RVpd), a Vision Transformer (ViT) with Nyström-based attention, a frozen-encoder
GRU predictor (RV-PL), and a combination of the former so that the GRU receives the
sequence of the embeddings the ViT produced. Event boundaries are detected by applying
peak detection to a change-score time series. Results are compared with those obtained
by an established baseline model (BL).
The approaches were tested on two datasets with distinct application scenarios: As-
sembly 101, a multi-view video-recording dataset featuring 53 participants assembling and
disassembling 101 toy vehicles, and Breakfast, a dataset of multi-view video recordings
of cooking activities. MCC with a ±15-frame tolerance was adopted as the primary per-
formance metric. On Assembly 101, results show MCCs of 0.405 (RVpd), 0.390 (EST),
0.228 (RV-PL, tying the baseline) and 0.228 (BL). For Breakfast, the results show MCCs
of 0.593 (RVpd), 0.514 (EST), 0.425 (RV-PL), and 0.256 (BL). A cross-corpus experi-
ment was also conducted by training the models on Assembly 101 and testing them on
Breakfast, yielding MCCs of 0.577 (RVpd), 0.522 (EST grid), and 0.425 (RV-PL).
A deeper investigation of the results reveals that the per-frame change score the mod-
els emit does carry boundary information (ground-truth frames score on average 1.75
times higher than | en_UK |
| dc.language.iso | en | |
| dc.language.iso | en | |
| dc.rights | info:eu-repo/semantics/openAccess | |
| dc.title | An EST-based Generic Event Boundary Detector | it_IT |
| dc.title.alternative | An EST-based Generic Event Boundary Detector | en_UK |
| dc.type | info:eu-repo/semantics/masterThesis | |
| dc.subject.miur | ING-INF/05 - SISTEMI DI ELABORAZIONE DELLE INFORMAZIONI | |
| dc.publisher.name | Università degli studi di Genova | |
| dc.date.academicyear | 2025/2026 | |
| dc.description.corsolaurea | 11661 - DIGITAL HUMANITIES - INTERACTIVE SYSTEMS AND DIGITAL MEDIA | |
| dc.description.area | 9 - INGEGNERIA | |
| dc.description.department | 100023 - DIPARTIMENTO DI INFORMATICA, BIOINGEGNERIA, ROBOTICA E INGEGNERIA DEI SISTEMI | |