# Fujisaki model

> Mediated Wiki article. Canonical URL: https://mediated.wiki/source/Fujisaki_model
> Markdown URL: https://mediated.wiki/source/Fujisaki_model.md
> Source: https://en.wikipedia.org/wiki/Fujisaki_model
> Source revision: 1322731065
> License: Creative Commons Attribution-ShareAlike 4.0 International (https://creativecommons.org/licenses/by-sa/4.0/)

{{Short description|Superpositional model}}
{{Technical|date=November 2020}}
thumb|An F0 contour is obtained by adding the phrase and accent components to the base frequency
The '''Fujisaki model''' is a [superpositional](/source/Superposition_principle) model for representing F<sub>0</sub> contour of [speech](/source/speech). 

According to the model, F<sub>0</sub> contour is generated as a result of the superposition of the outputs of two second order [linear filters](/source/Linear_filter) with a base frequency value. The second order linear filters are for generating the phrase and accent components of speech. The base frequency is the minimum frequency value of the speaker. In other words, F<sub>0</sub> contour is obtained by adding base frequency, phrase components and accent components. The model was proposed by Hiroya Fujisaki.

<math>
\ln(F_0(t)) = \ln(F_b) + \sum_{i=1}^I A_{pi} G_{pi} (t-T_{0i}) + \sum_{j=1}^J A_{aj} \{ G_{aj} (t-T_{1j}) - G_{aj} (t-T_{2j}) \} 
</math>
<br />
where
<br />
<math>
G_{pi}(t) = \alpha_i^2t \, \exp(-\alpha_i t) \quad \forall t \geq 0 ; = 0  \forall  t  \leq 0
</math>
<br />
<math display="inline">
G_{aj}(t) = \min[\gamma_j, \, 1-(1+\beta_j t) \, \exp(-\beta_j t)] \quad \forall t \geq 0 ; = 0  \forall  t  \leq 0
</math>

Where,

<math>F_b </math>: bias level upon which all the phrase and accent components are superposed to form an <math>F_0 </math> contour,

<math>I </math> : number of phrase commands,

<math>J </math> : number of accent commands,

<math>A_{pi} </math> : magnitude of the ith phrase command,

<math>A_{aj} </math> : amplitude of the jth accent command,

<math>T_{0i} </math> : instant of occurrence of the ith phrase
command,

<math>T_{1j} </math> : onset of the jth accent command,

<math>T_{2j} </math> : end of the jth accent command,

<math> \alpha_i </math> : natural angular frequency of the phrase control mechanism to the ith phrase command,

<math> \beta_j </math> : natural angular frequency of the accent control mechanism to the jth accent command, and

<math>\gamma_j </math> : ceiling level of the accent component for the jth accent command.

== References ==
*An Introduction to Text-to-Speech Synthesis<ref name="Dutoit2001">{{cite book|last=Dutoit|first=Thierry|title=An Introduction to Text-to-Speech Synthesis|year=2001|publisher=Kluwer Academic Publishers|isbn=1-4020-0369-2}}</ref> 
*{{Cite journal
 |author1=Keikichi Hirose |author2=Hiroya Fujisaki |author3=Mikio Yamaguchi | year = 1984
 | title = Synthesis by rule of voice fundamental frequency contours of spoken Japanese from linguistic information
 | journal = IEEE
 }}
{{reflist}}

Category:Speech
Category:Speech processing

---
Adapted from the Wikipedia article [Fujisaki model](https://en.wikipedia.org/wiki/Fujisaki_model) by Wikipedia contributors ([contributor history](https://en.wikipedia.org/wiki/Fujisaki_model?action=history)). Available under [Creative Commons Attribution-ShareAlike 4.0 International](https://creativecommons.org/licenses/by-sa/4.0/). Changes may have been made.
