Show simple item record

dc.contributor.advisorVolpe, Gualtiero <1974>
dc.contributor.advisorCeccaldi, Eleonora <1990>
dc.contributor.advisorVarni, Giovanna <1979>
dc.contributor.authorSerajoddin Mirghaed, Amirarsalan <1998>
dc.date.accessioned2026-07-23T14:32:21Z
dc.date.available2026-07-23T14:32:21Z
dc.date.issued2026-07-16
dc.identifier.urihttps://unire.unige.it/handle/123456789/16495
dc.description.abstractAbstract This thesis proposes a computational approach to Generic Event Boundary Detection (GEBD). GEBD involves identifying when meaningful events occur in an untrimmed video, without predefined action classes. The proposed approach is grounded in the Event Segmentation Theory (EST), which posits that humans detect boundaries by predicting errors when their internal model of the ongoing situation fails. EST is operationalised through a 15-channel feature representation aligned to EST perceptual categories of time, space, objects, characters, and interactions. Such a fea- ture representation is fed to three different machine-learning architectures: ReconViT-PD (RVpd), a Vision Transformer (ViT) with Nyström-based attention, a frozen-encoder GRU predictor (RV-PL), and a combination of the former so that the GRU receives the sequence of the embeddings the ViT produced. Event boundaries are detected by applying peak detection to a change-score time series. Results are compared with those obtained by an established baseline model (BL). The approaches were tested on two datasets with distinct application scenarios: As- sembly 101, a multi-view video-recording dataset featuring 53 participants assembling and disassembling 101 toy vehicles, and Breakfast, a dataset of multi-view video recordings of cooking activities. MCC with a ±15-frame tolerance was adopted as the primary per- formance metric. On Assembly 101, results show MCCs of 0.405 (RVpd), 0.390 (EST), 0.228 (RV-PL, tying the baseline) and 0.228 (BL). For Breakfast, the results show MCCs of 0.593 (RVpd), 0.514 (EST), 0.425 (RV-PL), and 0.256 (BL). A cross-corpus experi- ment was also conducted by training the models on Assembly 101 and testing them on Breakfast, yielding MCCs of 0.577 (RVpd), 0.522 (EST grid), and 0.425 (RV-PL). A deeper investigation of the results reveals that the per-frame change score the mod- els emit does carry boundary information (ground-truth frames score on average 1.75 times higher thanit_IT
dc.description.abstractAbstract This thesis proposes a computational approach to Generic Event Boundary Detection (GEBD). GEBD involves identifying when meaningful events occur in an untrimmed video, without predefined action classes. The proposed approach is grounded in the Event Segmentation Theory (EST), which posits that humans detect boundaries by predicting errors when their internal model of the ongoing situation fails. EST is operationalised through a 15-channel feature representation aligned to EST perceptual categories of time, space, objects, characters, and interactions. Such a fea- ture representation is fed to three different machine-learning architectures: ReconViT-PD (RVpd), a Vision Transformer (ViT) with Nyström-based attention, a frozen-encoder GRU predictor (RV-PL), and a combination of the former so that the GRU receives the sequence of the embeddings the ViT produced. Event boundaries are detected by applying peak detection to a change-score time series. Results are compared with those obtained by an established baseline model (BL). The approaches were tested on two datasets with distinct application scenarios: As- sembly 101, a multi-view video-recording dataset featuring 53 participants assembling and disassembling 101 toy vehicles, and Breakfast, a dataset of multi-view video recordings of cooking activities. MCC with a ±15-frame tolerance was adopted as the primary per- formance metric. On Assembly 101, results show MCCs of 0.405 (RVpd), 0.390 (EST), 0.228 (RV-PL, tying the baseline) and 0.228 (BL). For Breakfast, the results show MCCs of 0.593 (RVpd), 0.514 (EST), 0.425 (RV-PL), and 0.256 (BL). A cross-corpus experi- ment was also conducted by training the models on Assembly 101 and testing them on Breakfast, yielding MCCs of 0.577 (RVpd), 0.522 (EST grid), and 0.425 (RV-PL). A deeper investigation of the results reveals that the per-frame change score the mod- els emit does carry boundary information (ground-truth frames score on average 1.75 times higher thanen_UK
dc.language.isoen
dc.language.isoen
dc.rightsinfo:eu-repo/semantics/openAccess
dc.titleAn EST-based Generic Event Boundary Detectorit_IT
dc.title.alternativeAn EST-based Generic Event Boundary Detectoren_UK
dc.typeinfo:eu-repo/semantics/masterThesis
dc.subject.miurING-INF/05 - SISTEMI DI ELABORAZIONE DELLE INFORMAZIONI
dc.publisher.nameUniversità degli studi di Genova
dc.date.academicyear2025/2026
dc.description.corsolaurea11661 - DIGITAL HUMANITIES - INTERACTIVE SYSTEMS AND DIGITAL MEDIA
dc.description.area9 - INGEGNERIA
dc.description.department100023 - DIPARTIMENTO DI INFORMATICA, BIOINGEGNERIA, ROBOTICA E INGEGNERIA DEI SISTEMI


Files in this item

Thumbnail

This item appears in the following Collection(s)

Show simple item record