Deep End-to-End Representation Learning for Food Type Recognition from Speech
| dc.contributor.author | Sertolli, Benjamin | |
| dc.contributor.author | Cummins, Nicholas | |
| dc.contributor.author | Sengur, Abdulkadir | |
| dc.contributor.author | Schuller, Bjorn W. | |
| dc.date.accessioned | 2026-08-12T16:41:35Z | |
| dc.date.issued | 2018 | |
| dc.department | Fırat Üniversitesi | |
| dc.description | 20th ACM International Conference on Multimodal Interaction (ICMI) -- OCT 16-20, 2018 -- Boulder, CO | |
| dc.description.abstract | The use of Convolutional Neural Networks (CNN) pre-trained for a particular task, as a feature extractor for an alternate task, is a standard practice in many image classification paradigms. However, to date there have been comparatively few works exploring this technique for speech classification tasks. Herein, we utilise a pre-trained end-to-end Automatic Speech Recognition CNN as a feature extractor for the task of food-type recognition from speech. Furthermore, we also explore the benefits of Compact Bilinear Pooling for combining multiple feature representations extracted from the CNN. Key results presented indicate the suitability of this approach. When combined with a Recurrent Neural Network classifier, our strongest system achieves, for a seven-class food-type classification task an unweighted average recall of 73.3 % on the test set of the IHEARu-EAT database. | |
| dc.description.sponsorship | European Unions [338164] | |
| dc.description.sponsorship | This work was supported by the European Unions's Seventh Framework and Horizon 2020 Programmes under grant agreement No. 338164 (ERC StG iHEARu). | |
| dc.description.sponsorship | Assoc Comp Machinery,Assoc Comp Machinery SIGCHI,Openstream,Microsoft,Univ Colorado Boulder, Inst Cognit Sci,audEERING | |
| dc.identifier.doi | 10.1145/3242969.3243683 | |
| dc.identifier.endpage | 578 | |
| dc.identifier.isbn | 978-1-4503-5692-3 | |
| dc.identifier.orcid | 0000-0002-6478-8699 | |
| dc.identifier.orcid | 0000-0002-1178-917X | |
| dc.identifier.orcid | 0000-0003-1614-2639 | |
| dc.identifier.scopus | 2-s2.0-85056613029 | |
| dc.identifier.scopusquality | N/A | |
| dc.identifier.startpage | 574 | |
| dc.identifier.uri | https://doi.org/10.1145/3242969.3243683 | |
| dc.identifier.uri | https://hdl.handle.net/11508/45898 | |
| dc.identifier.wos | WOS:000457913100087 | |
| dc.identifier.wosquality | N/A | |
| dc.indekslendigikaynak | Web of Science | |
| dc.indekslendigikaynak | Scopus | |
| dc.language.iso | en | |
| dc.publisher | Assoc Computing Machinery | |
| dc.relation.ispartof | Icmi'18: Proceedings of the 20Th Acm International Conference on Multimodal Interaction | |
| dc.relation.publicationcategory | Konferans Öğesi - Uluslararası - Kurum Öğretim Elemanı | |
| dc.rights | info:eu-repo/semantics/openAccess | |
| dc.snmz | KA_WoS_20260511 | |
| dc.subject | Eating Condition | |
| dc.subject | Deep Representation Learning | |
| dc.subject | End-to-End Learning | |
| dc.subject | Compact Bilinear Pooling | |
| dc.subject | Recurrent Neural Networks | |
| dc.title | Deep End-to-End Representation Learning for Food Type Recognition from Speech | |
| dc.type | Conference Object |







