Deep End-to-End Representation Learning for Food Type Recognition from Speech

dc.contributor.authorSertolli, Benjamin
dc.contributor.authorCummins, Nicholas
dc.contributor.authorSengur, Abdulkadir
dc.contributor.authorSchuller, Bjorn W.
dc.date.accessioned2026-08-12T16:41:35Z
dc.date.issued2018
dc.departmentFırat Üniversitesi
dc.description20th ACM International Conference on Multimodal Interaction (ICMI) -- OCT 16-20, 2018 -- Boulder, CO
dc.description.abstractThe use of Convolutional Neural Networks (CNN) pre-trained for a particular task, as a feature extractor for an alternate task, is a standard practice in many image classification paradigms. However, to date there have been comparatively few works exploring this technique for speech classification tasks. Herein, we utilise a pre-trained end-to-end Automatic Speech Recognition CNN as a feature extractor for the task of food-type recognition from speech. Furthermore, we also explore the benefits of Compact Bilinear Pooling for combining multiple feature representations extracted from the CNN. Key results presented indicate the suitability of this approach. When combined with a Recurrent Neural Network classifier, our strongest system achieves, for a seven-class food-type classification task an unweighted average recall of 73.3 % on the test set of the IHEARu-EAT database.
dc.description.sponsorshipEuropean Unions [338164]
dc.description.sponsorshipThis work was supported by the European Unions's Seventh Framework and Horizon 2020 Programmes under grant agreement No. 338164 (ERC StG iHEARu).
dc.description.sponsorshipAssoc Comp Machinery,Assoc Comp Machinery SIGCHI,Openstream,Microsoft,Univ Colorado Boulder, Inst Cognit Sci,audEERING
dc.identifier.doi10.1145/3242969.3243683
dc.identifier.endpage578
dc.identifier.isbn978-1-4503-5692-3
dc.identifier.orcid0000-0002-6478-8699
dc.identifier.orcid0000-0002-1178-917X
dc.identifier.orcid0000-0003-1614-2639
dc.identifier.scopus2-s2.0-85056613029
dc.identifier.scopusqualityN/A
dc.identifier.startpage574
dc.identifier.urihttps://doi.org/10.1145/3242969.3243683
dc.identifier.urihttps://hdl.handle.net/11508/45898
dc.identifier.wosWOS:000457913100087
dc.identifier.wosqualityN/A
dc.indekslendigikaynakWeb of Science
dc.indekslendigikaynakScopus
dc.language.isoen
dc.publisherAssoc Computing Machinery
dc.relation.ispartofIcmi'18: Proceedings of the 20Th Acm International Conference on Multimodal Interaction
dc.relation.publicationcategoryKonferans Öğesi - Uluslararası - Kurum Öğretim Elemanı
dc.rightsinfo:eu-repo/semantics/openAccess
dc.snmzKA_WoS_20260511
dc.subjectEating Condition
dc.subjectDeep Representation Learning
dc.subjectEnd-to-End Learning
dc.subjectCompact Bilinear Pooling
dc.subjectRecurrent Neural Networks
dc.titleDeep End-to-End Representation Learning for Food Type Recognition from Speech
dc.typeConference Object

Dosyalar