There is provided a computer-implemented method of training a speech-to-speech (S2S) machine learning (ML) model for adapting voice attribute(s) of speech, in which the duration of phones of the second audio content are controlled in response to segment-level durations defined by segment-level start and end time stamps.