Università di Genova logo, link al sitoUniRe logo, link alla pagina iniziale
    • English
    • italiano
  • italiano 
    • English
    • italiano
  • Login
Mostra Item 
  •   Home
  • Tesi
  • Tesi di Laurea
  • Laurea Magistrale
  • Mostra Item
  •   Home
  • Tesi
  • Tesi di Laurea
  • Laurea Magistrale
  • Mostra Item
JavaScript is disabled for your browser. Some features of this site may not work without it.

An EST-based Generic Event Boundary Detector

Thumbnail
Mostra/Apri
tesi38024964.pdf (1.021Mb)
Autore
Serajoddin Mirghaed, Amirarsalan <1998>
Data
2026-07-16
Disponibile dal
2026-07-23
Abstract
Abstract This thesis proposes a computational approach to Generic Event Boundary Detection (GEBD). GEBD involves identifying when meaningful events occur in an untrimmed video, without predefined action classes. The proposed approach is grounded in the Event Segmentation Theory (EST), which posits that humans detect boundaries by predicting errors when their internal model of the ongoing situation fails. EST is operationalised through a 15-channel feature representation aligned to EST perceptual categories of time, space, objects, characters, and interactions. Such a fea- ture representation is fed to three different machine-learning architectures: ReconViT-PD (RVpd), a Vision Transformer (ViT) with Nyström-based attention, a frozen-encoder GRU predictor (RV-PL), and a combination of the former so that the GRU receives the sequence of the embeddings the ViT produced. Event boundaries are detected by applying peak detection to a change-score time series. Results are compared with those obtained by an established baseline model (BL). The approaches were tested on two datasets with distinct application scenarios: As- sembly 101, a multi-view video-recording dataset featuring 53 participants assembling and disassembling 101 toy vehicles, and Breakfast, a dataset of multi-view video recordings of cooking activities. MCC with a ±15-frame tolerance was adopted as the primary per- formance metric. On Assembly 101, results show MCCs of 0.405 (RVpd), 0.390 (EST), 0.228 (RV-PL, tying the baseline) and 0.228 (BL). For Breakfast, the results show MCCs of 0.593 (RVpd), 0.514 (EST), 0.425 (RV-PL), and 0.256 (BL). A cross-corpus experi- ment was also conducted by training the models on Assembly 101 and testing them on Breakfast, yielding MCCs of 0.577 (RVpd), 0.522 (EST grid), and 0.425 (RV-PL). A deeper investigation of the results reveals that the per-frame change score the mod- els emit does carry boundary information (ground-truth frames score on average 1.75 times higher than
 
Abstract This thesis proposes a computational approach to Generic Event Boundary Detection (GEBD). GEBD involves identifying when meaningful events occur in an untrimmed video, without predefined action classes. The proposed approach is grounded in the Event Segmentation Theory (EST), which posits that humans detect boundaries by predicting errors when their internal model of the ongoing situation fails. EST is operationalised through a 15-channel feature representation aligned to EST perceptual categories of time, space, objects, characters, and interactions. Such a fea- ture representation is fed to three different machine-learning architectures: ReconViT-PD (RVpd), a Vision Transformer (ViT) with Nyström-based attention, a frozen-encoder GRU predictor (RV-PL), and a combination of the former so that the GRU receives the sequence of the embeddings the ViT produced. Event boundaries are detected by applying peak detection to a change-score time series. Results are compared with those obtained by an established baseline model (BL). The approaches were tested on two datasets with distinct application scenarios: As- sembly 101, a multi-view video-recording dataset featuring 53 participants assembling and disassembling 101 toy vehicles, and Breakfast, a dataset of multi-view video recordings of cooking activities. MCC with a ±15-frame tolerance was adopted as the primary per- formance metric. On Assembly 101, results show MCCs of 0.405 (RVpd), 0.390 (EST), 0.228 (RV-PL, tying the baseline) and 0.228 (BL). For Breakfast, the results show MCCs of 0.593 (RVpd), 0.514 (EST), 0.425 (RV-PL), and 0.256 (BL). A cross-corpus experi- ment was also conducted by training the models on Assembly 101 and testing them on Breakfast, yielding MCCs of 0.577 (RVpd), 0.522 (EST grid), and 0.425 (RV-PL). A deeper investigation of the results reveals that the per-frame change score the mod- els emit does carry boundary information (ground-truth frames score on average 1.75 times higher than
 
Tipo
info:eu-repo/semantics/masterThesis
Collezioni
  • Laurea Magistrale [8032]
URI
https://unire.unige.it/handle/123456789/16495
Metadati
Mostra tutti i dati dell'item

UniRe - Università degli studi di Genova | Informazioni e Supporto
 

 

UniReArchivi & Collezioni

Area personale

Login

UniRe - Università degli studi di Genova | Informazioni e Supporto