In this paper, we argue that multiword expression identif ication systems based on BERT are able to capture semi-productive patterns that generate multiword expressions. To test this hypothesis we analyzed the results obtained by MTLB-STRUCT on unseen multiword expressions during edition 1.2 of the PARSEME shared task. We observed that MTLB-STRUCT discovers, in proportion, more light verb constructions and verb particle constructions than verbal idioms. Since light verb constructions and verb-particle constructions often result from semi-productive patterns, while verbal idioms are more idiosyncratic, the results corroborate our hypothesis.
The Role of Semi-productivity in Multiword Expression Identification: Why can BERT Capture novel MWEs?
Cirillo, Nicola
;Paone, Antonietta
2022-01-01
Abstract
In this paper, we argue that multiword expression identif ication systems based on BERT are able to capture semi-productive patterns that generate multiword expressions. To test this hypothesis we analyzed the results obtained by MTLB-STRUCT on unseen multiword expressions during edition 1.2 of the PARSEME shared task. We observed that MTLB-STRUCT discovers, in proportion, more light verb constructions and verb particle constructions than verbal idioms. Since light verb constructions and verb-particle constructions often result from semi-productive patterns, while verbal idioms are more idiosyncratic, the results corroborate our hypothesis.I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.