<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"
	xmlns:content="http://purl.org/rss/1.0/modules/content/"
	xmlns:wfw="http://wellformedweb.org/CommentAPI/"
	xmlns:dc="http://purl.org/dc/elements/1.1/"
	xmlns:atom="http://www.w3.org/2005/Atom"
	xmlns:sy="http://purl.org/rss/1.0/modules/syndication/"
	xmlns:slash="http://purl.org/rss/1.0/modules/slash/"
	>

<channel>
	<title>Jhonatan López</title>
	<atom:link href="https://www.jhonatanlopez.com/feed/" rel="self" type="application/rss+xml" />
	<link>https://www.jhonatanlopez.com/</link>
	<description>Engineering &#38; Sound Design</description>
	<lastBuildDate>Thu, 20 Aug 2026 17:05:02 +0000</lastBuildDate>
	<language>en-GB</language>
	<sy:updatePeriod>
	hourly	</sy:updatePeriod>
	<sy:updateFrequency>
	1	</sy:updateFrequency>
	

<image>
	<url>https://www.jhonatanlopez.com/wp-content/uploads/2019/01/cropped-Logo-Web-Jhonatan2-1-32x32.png</url>
	<title>Jhonatan López</title>
	<link>https://www.jhonatanlopez.com/</link>
	<width>32</width>
	<height>32</height>
</image> 
	<item>
		<title>Andes JX: A Synthesizer from Latitude Zero</title>
		<link>https://www.jhonatanlopez.com/andean-synthesizer-andes-jx/</link>
					<comments>https://www.jhonatanlopez.com/andean-synthesizer-andes-jx/#comments</comments>
		
		<dc:creator><![CDATA[Jhonatan López]]></dc:creator>
		<pubDate>Thu, 30 Jul 2026 22:06:04 +0000</pubDate>
				<category><![CDATA[Audio Software Development]]></category>
		<category><![CDATA[Research]]></category>
		<category><![CDATA[Sound Design]]></category>
		<category><![CDATA[AES LAC 2026]]></category>
		<category><![CDATA[Andes JX]]></category>
		<category><![CDATA[audio DSP]]></category>
		<category><![CDATA[Ecuador]]></category>
		<category><![CDATA[JUCE]]></category>
		<category><![CDATA[open source]]></category>
		<category><![CDATA[synthesizer]]></category>
		<guid isPermaLink="false">https://www.jhonatanlopez.com/?p=5454</guid>

					<description><![CDATA[<p>Subtractive synthesis, Andean identity, and open audio software from Ecuador. Andes JX user interface. Its slate-blue palette and restrained visual language draw inspiration from stone, sky, and the high-altitude landscapes of the Ecuadorian Andes. A polysynth from Ecuador There is no completely neutral place from which to design an instrument. Every synthesizer reflects technical decisions, [&#8230;]</p>
<p>The post <a href="https://www.jhonatanlopez.com/andean-synthesizer-andes-jx/">Andes JX: A Synthesizer from Latitude Zero</a> appeared first on <a href="https://www.jhonatanlopez.com">Jhonatan López</a>.</p>
]]></description>
										<content:encoded><![CDATA[
<p class="wp-block-paragraph"><em>Subtractive synthesis, Andean identity, and open audio software from Ecuador.</em></p>



<figure class="wp-block-image size-large">
  <img decoding="async"
    src="https://www.jhonatanlopez.com/wp-content/uploads/2026/07/ScreenshotAndesJX2.png"
    alt="Andes JX open-source Andean synthesizer user interface"
    class="wp-image-1850"
  >
  <figcaption>
    Andes JX user interface. Its slate-blue palette and restrained visual language draw inspiration from stone, sky, and the high-altitude landscapes of the Ecuadorian Andes.
  </figcaption>
</figure>



<h2 class="wp-block-heading">A polysynth from Ecuador</h2>



<p class="wp-block-paragraph">There is no completely neutral place from which to design an instrument. Every synthesizer reflects technical decisions, visual references, and listening habits shaped by the people and environments behind it.</p>



<p class="wp-block-paragraph">Instruments created in the United States, Europe, and Japan shaped much of the vocabulary of modern subtractive synthesis. Andes JX builds on that legacy while articulating it from a different territory: Quito, Ecuador.</p>



<p class="wp-block-paragraph">Andes JX is an open-source polyphonic subtractive synthesizer developed in C++ with the JUCE framework. It was conceived as a musical instrument, an open software project, and a platform for studying digital signal processing and audio software architecture.</p>



<p class="wp-block-paragraph">In 2026, Andes JX was presented at AES LAC 2026 in Lima, Peru, as part of the convention’s technical program.</p>



<p class="wp-block-paragraph">The project does not attempt to invent a new synthesis paradigm. Instead, it asks what happens when a familiar architecture is designed from a different geography, visual memory, and cultural perspective.</p>



<h2 class="wp-block-heading">A familiar architecture from a different place</h2>



<p class="wp-block-paragraph">From a technical perspective, Andes JX uses the fundamental building blocks of subtractive synthesis: oscillators that generate harmonically rich signals, filters that reshape their spectrum, envelopes that control their evolution over time, and a low-frequency oscillator for modulation.</p>



<p class="wp-block-paragraph">In addition, the instrument includes two oscillators with PolyBLEP-based aliasing reduction, independent amplitude and filter envelopes, global LFO modulation, multiple polyphony modes, glide, keyboard tracking, velocity sensitivity, and MIDI control.</p>



<p class="wp-block-paragraph">One of its most significant sonic decisions is the inclusion of two switchable filter models.</p>



<p class="wp-block-paragraph">The first is a State Variable Filter based on a topology-preserving structure. It provides a stable and controlled response while supporting several filtering modes within a flexible architecture.</p>



<p class="wp-block-paragraph">By contrast, the nonlinear Moog Ladder filter introduces stronger saturation and a more pronounced resonant character.</p>



<p class="wp-block-paragraph">Together, these two filters reflect one of the instrument’s central design principles: keep the architecture focused, but ensure that each decision produces an audible result. Andes JX does not aim to provide unlimited modules or endless modulation routing. Instead, its interface offers a deliberate set of controls, each with a clearly defined role.</p>



<h2 class="wp-block-heading">Andean geography as identity</h2>



<p class="wp-block-paragraph">The Andean identity of Andes JX is not intended to reproduce Ecuadorian traditional music literally or to translate a landscape directly into sound. Instead, it appears through visual, conceptual, and sonic decisions that place the instrument within a specific territory.</p>



<h3 class="wp-block-heading">Visual language</h3>



<p class="wp-block-paragraph">The interface uses a slate-blue palette, soft off-white text, and a restrained composition inspired by stone, mist, sky, and the high-altitude landscapes of the Ecuadorian Andes.</p>



<p class="wp-block-paragraph">Rather than relying on bright accent colors or a futuristic aesthetic, the instrument adopts a contained visual language intended to suggest openness, stability, and the severity of mountain environments.</p>



<h3 class="wp-block-heading">Factory presets</h3>



<p class="wp-block-paragraph">The 33 factory presets are named after real locations in the Ecuadorian Andes, combined with musical or timbral references.</p>



<p class="wp-block-paragraph">Examples include:</p>



<ul class="wp-block-list">
<li>Cotopaxi Acid;</li>



<li>Cayambe 5th;</li>



<li>Imbabura Pedal;</li>



<li>Quilotoa Aqua;</li>



<li>Páramo Sostenuto.</li>
</ul>



<p class="wp-block-paragraph">These names propose associations between geography and sonic character. They are not intended as literal representations of each place. Instead, the names, landscapes, and cultural associations provide starting points for sound design.</p>



<p class="wp-block-paragraph">The bank retains familiar subtractive-synthesis categories such as basses, pads, leads, brass, keys, plucks, and effects. However, it presents them through an identity connected to Ecuadorian geography.</p>



<h3 class="wp-block-heading">An open educational project</h3>



<p class="wp-block-paragraph">Andes JX was not developed only as a production tool. It was also designed to show how a complete polyphonic synthesizer is organized internally.</p>



<p class="wp-block-paragraph">The source code is available under the GNU GPL v3 license, together with bilingual documentation in English and Spanish. Its architecture separates the main responsibilities of the system:</p>



<ul class="wp-block-list">
<li>DSP processing;</li>



<li>plugin, parameter, and MIDI management;</li>



<li>preset handling;</li>



<li>graphical interface;</li>



<li>host integration.</li>
</ul>



<p class="wp-block-paragraph">This separation allows individual modules to be studied independently while also showing how they interact within a real-time audio application.</p>



<p class="wp-block-paragraph">Technical and academic references accompany the PolyBLEP oscillators, Moog Ladder filter, and State Variable Filter implementations. In this way, a student or developer can begin with the source code, consult the original publications, and follow the path from a mathematical concept to its application within a functional instrument.</p>



<p class="wp-block-paragraph">The project does not claim to present the only correct synthesizer architecture. Its educational value lies in offering a complete implementation that can be read, tested, questioned, and modified.</p>



<h3 class="wp-block-heading">The JX lineage</h3>



<p class="wp-block-paragraph">The name Andes JX acknowledges two lines of influence.</p>



<p class="wp-block-paragraph">The first comes from the Roland JX family of polyphonic synthesizers, known for their subtractive architecture and their association with the analog instruments of the 1980s.</p>



<p class="wp-block-paragraph">The second belongs to open and educational software.</p>



<p class="wp-block-paragraph">Andes JX grew from the study of JX11, the educational synthesizer developed in JUCE by Matthijs Hollemans for <em>Creating Synthesizer Plug-Ins with C++ and JUCE</em>. JX11, in turn, is related to MDA JX10, the freeware synthesizer created by Paul Kellett and later released as open-source software.</p>



<p class="wp-block-paragraph">Andes JX continues that lineage while expanding the architecture, documentation, visual identity, and academic purpose of the project.</p>



<p class="wp-block-paragraph">The term “JX” therefore acknowledges its technical ancestry, while “Andes” defines the place from which that lineage continues.</p>



<h2 class="wp-block-heading">From NoiseRoomUIO to AES LAC 2026</h2>



<p class="wp-block-paragraph">Andes JX is part of the NoiseRoom UIO project catalog, alongside Classic Compressor CCMkI and NoiseRverb.</p>



<p class="wp-block-paragraph">Although each plugin addresses a different technical problem, all three share a common intention: combining audio software development with an identity connected to the context in which the projects were created.</p>



<p class="wp-block-paragraph">NoiseRverb used impulse responses captured in architectural spaces in Quito. Andes JX extends that search into synthesis and virtual instrument design.</p>



<p class="wp-block-paragraph">The development of the synthesizer also led to the academic paper:</p>



<p class="wp-block-paragraph"><strong>“Design and Architecture of a Polyphonic Subtractive Synthesizer as a Modular Structure for Teaching DSP and Audio Software.”</strong></p>



<p class="wp-block-paragraph">The paper was selected by the AES LAC 2026 review committee and presented at the convention in Lima, Peru. It examines how a functional synthesizer can serve as an object of study for teaching digital signal processing and audio software engineering.</p>



<p class="wp-block-paragraph">Presenting Andes JX at AES LAC 2026 provided an opportunity to share work developed in Ecuador with a Latin American community of researchers, educators, engineers, and audio developers.</p>



<h3 class="wp-block-heading">Try Andes JX</h3>



<p class="wp-block-paragraph">You can download Andes JX as a VST3 plugin for Windows.</p>



<p class="wp-block-paragraph">The project includes:</p>



<ul class="wp-block-list">
<li>33 factory presets;</li>



<li>C++ source code;</li>



<li>a bilingual user manual;</li>



<li>technical documentation in English and Spanish;</li>



<li>a modular architecture built with JUCE;</li>



<li>GNU GPL v3 licensing.</li>
</ul>



<p class="wp-block-paragraph">You can download the instrument, explore the source code, and review the documentation through the following links:</p>



<p class="wp-block-paragraph"><a href="https://www.jhonatanlopez.com/wp-content/uploads/2026/07/Andes-JX-Setup-v1.0.0.zip"><strong>Download Andes JX</strong></a><br><a href="https://github.com/bansky0/Andes-JX"><strong>GitHub repository</strong></a><br><a href="https://github.com/bansky0/Andes-JX/blob/main/Docs/manualAndesJxESP.md"><strong>User manual</strong></a></p>



<p class="wp-block-paragraph">Andes JX is an Andean synthesizer built in Ecuador, inspired by the Andes, and designed to be heard, studied, and transformed.</p>



<p class="wp-block-paragraph"><strong>Development:</strong> NoiseRoomUIO, Quito, Ecuador<br><strong>Authors:</strong> Jhonatan López and Valeria Villarreal<br><strong>Academic context:</strong> work developed in connection with the Numerical Simulation and Computational Analysis Laboratory at Escuela Politécnica Nacional</p>



<p class="wp-block-paragraph"></p>
<p>The post <a href="https://www.jhonatanlopez.com/andean-synthesizer-andes-jx/">Andes JX: A Synthesizer from Latitude Zero</a> appeared first on <a href="https://www.jhonatanlopez.com">Jhonatan López</a>.</p>
]]></content:encoded>
					
					<wfw:commentRss>https://www.jhonatanlopez.com/andean-synthesizer-andes-jx/feed/</wfw:commentRss>
			<slash:comments>1</slash:comments>
		
		
			</item>
		<item>
		<title>NoiseRverb: Heritage Acoustics in a Plugin</title>
		<link>https://www.jhonatanlopez.com/heritage-acoustics-plugin/</link>
		
		<dc:creator><![CDATA[Jhonatan López]]></dc:creator>
		<pubDate>Sat, 31 May 2025 02:00:00 +0000</pubDate>
				<category><![CDATA[DSP]]></category>
		<category><![CDATA[Plugin]]></category>
		<guid isPermaLink="false">https://www.jhonatanlopez.com/?p=4853</guid>

					<description><![CDATA[<p>New blog post — I’ve been meaning to write this for a long time. It’s about a concept that has fascinated me from the moment I encountered it. I’ll also share a small contribution I made in this area with NoiseRverb, a heritage acoustics plugin that captures the sonic identity of historic churches in Quito [&#8230;]</p>
<p>The post <a href="https://www.jhonatanlopez.com/heritage-acoustics-plugin/">NoiseRverb: Heritage Acoustics in a Plugin</a> appeared first on <a href="https://www.jhonatanlopez.com">Jhonatan López</a>.</p>
]]></description>
										<content:encoded><![CDATA[
<figure class="wp-block-image size-large"><img fetchpriority="high" decoding="async" width="1024" height="471" src="https://www.jhonatanlopez.com/wp-content/uploads/2025/05/image-1024x471.png" alt="NoiseRverb heritage acoustics plugin interface." class="wp-image-4854" srcset="https://www.jhonatanlopez.com/wp-content/uploads/2025/05/image-1024x471.png 1024w, https://www.jhonatanlopez.com/wp-content/uploads/2025/05/image-600x276.png 600w, https://www.jhonatanlopez.com/wp-content/uploads/2025/05/image-300x138.png 300w, https://www.jhonatanlopez.com/wp-content/uploads/2025/05/image-768x353.png 768w, https://www.jhonatanlopez.com/wp-content/uploads/2025/05/image.png 1137w" sizes="(max-width: 1024px) 100vw, 1024px" /></figure>



<p class="wp-block-paragraph">New blog post — I’ve been meaning to write this for a long time. It’s about a concept that has fascinated me from the moment I encountered it. I’ll also share a small contribution I made in this area with <em>NoiseRverb</em>, a heritage acoustics plugin that captures the sonic identity of historic churches in Quito through convolution reverb.</p>



<p class="wp-block-paragraph">We often associate heritage with visual beauty—particularly in architecture. This connection largely stems from the preservation of artistic landmarks due to their historical significance. As a result, many of these sites have become popular tourist attractions.</p>



<p class="wp-block-paragraph">But what if we consider a different kind of heritage — one we can hear? Can sound hold historical value? Can it be preserved? The answer is yes. While this isn’t a new idea — researchers have studied it before — it’s starting to find its place in the audio industry as well.</p>



<p class="wp-block-paragraph">While some heritage spaces are preserved through sound art, others live on through live recordings or songs. For example, the <em>Festival Internacional de Música Sacra</em> (FIMUSAQ) in Quito has helped preserve local culture through sacred music. While it&#8217;s not the only example globally, it clearly shows how sound can carry historical weight.</p>



<h2 class="wp-block-heading">What Is Heritage Acoustics?</h2>



<p class="wp-block-paragraph">Heritage acoustics is a field that explores and protects the sound of historic spaces. It combines architecture, physics, and sound engineering. From cathedrals to theatres, their sound is as iconic as their structure.</p>



<p class="wp-block-paragraph">The goal is to study how sound behaves in these spaces. It involves looking at reverberation, reflections, absorption, and how sound waves move. There are three main techniques used:</p>



<ul class="wp-block-list">
<li><strong>On-site measurements</strong>: Microphones and speakers are placed in the space. Specific signals are played (like sine sweeps), and the room&#8217;s response is recorded. This captures real acoustic data.</li>



<li><strong>Impulse response (IR) capture</strong>: This technique records how a room reacts to a quick, broad sound. The result is a unique audio fingerprint of the space.</li>



<li><strong>Computer modelling</strong>: 3D simulations of the space are created using architectural drawings. This allows engineers to predict how sound will behave, even without access to the real location.</li>
</ul>



<p class="wp-block-paragraph">These methods help protect the unique sound of historical buildings and let others experience it — even from far away.</p>



<h2 class="wp-block-heading">Applications in the Music Industry</h2>



<p class="wp-block-paragraph">The study of heritage acoustics has led to exciting tools for musicians and producers. These include:</p>



<ul class="wp-block-list">
<li>Convolution reverb plugins that recreate historic spaces.</li>



<li>Acoustic modelling for virtual concerts.</li>



<li>Restoring old recordings with their original acoustics.</li>



<li>Immersive museum or VR audio experiences.</li>
</ul>



<p class="wp-block-paragraph">All of these are interesting, but I’ll focus on convolution reverb, as it’s where I work most.</p>



<h2 class="wp-block-heading">Convolution Reverb</h2>



<p class="wp-block-paragraph">To understand convolution reverb, you need to know what an <strong>impulse response (IR)</strong> is. An IR is a recording of how a space reacts to a short, full-range sound — like a clap or a burst of noise. It contains key information about the room’s reflections and decay.</p>



<p class="wp-block-paragraph">There are many ways to create an IR: clapping, popping a balloon, or playing a sine sweep. The results are similar, no matter the method.</p>



<p class="wp-block-paragraph">Convolution reverb takes this IR and applies it to any audio signal. The result is a realistic simulation of how that signal would sound in the original space. This is done using a mathematical process called convolution. It combines the original sound with the impulse response to recreate the acoustic experience.</p>



<p class="wp-block-paragraph">This technique is popular for its realism. It lets musicians take a studio recording and place it in a real, historic space — without leaving their DAW.</p>



<h2 class="wp-block-heading">Heritage Acoustics Plugin</h2>



<p class="wp-block-paragraph">As part of our exploration, I worked with engineers <strong>Analí Pinto</strong> and <strong>Fausto Espinoza</strong> to create a VST3 plugin called <em>NoiseRverb</em>. It’s based on impulse responses captured in seven churches in Quito: <strong>San Francisco, Basílica, Catedral, Compañía, Guápulo, El Sagrario, and Santo Domingo</strong>.</p>



<p class="wp-block-paragraph">The plugin lets musicians and producers experience these spaces inside their usual music-making software.</p>



<p class="wp-block-paragraph">We recorded the impulse responses on-site and processed them digitally. Then, we used real-time convolution to bring those spaces to life inside the plugin. The result is an authentic, immersive experience.</p>



<p class="wp-block-paragraph"><strong>NoiseRverb </strong>is free to download<strong> </strong><a href="https://www.jhonatanlopez.com/sound-design/">here</a><strong>.</strong></p>



<p class="wp-block-paragraph">Learn more about heritage acoustics and its role in cultural preservation in this <a class="" href="https://www.sciencedirect.com/topics/engineering/architectural-acoustics">introductory article on architectural acoustics</a> and in the <a class="" href="https://www.mdpi.com/2075-5309/15/15/2639">study on acoustic and perceptual variables in three heritage churches of Quito</a>.</p>



<p class="wp-block-paragraph"></p>
<p>The post <a href="https://www.jhonatanlopez.com/heritage-acoustics-plugin/">NoiseRverb: Heritage Acoustics in a Plugin</a> appeared first on <a href="https://www.jhonatanlopez.com">Jhonatan López</a>.</p>
]]></content:encoded>
					
		
		
			</item>
		<item>
		<title>AllYouNeedIsSound 3: Spectral Representations and Feature Extraction</title>
		<link>https://www.jhonatanlopez.com/advanced-spectral-representations-audio-analysis/</link>
		
		<dc:creator><![CDATA[Jhonatan López]]></dc:creator>
		<pubDate>Mon, 24 Mar 2025 13:33:36 +0000</pubDate>
				<category><![CDATA[DSP]]></category>
		<category><![CDATA[analysis]]></category>
		<guid isPermaLink="false">https://www.jhonatanlopez.com/?p=4602</guid>

					<description><![CDATA[<p>Have you ever wondered how machines understand the nuances of sound? In my previous post, we explored spectral analysis and learned how spectrograms reveal the frequency content of audio signals using the Short-Time Fourier Transform (STFT). Now, let’s dive deeper into advanced spectral representations for audio analysis, including Mel Spectrograms, CQT, and HCQT, and show how they [&#8230;]</p>
<p>The post <a href="https://www.jhonatanlopez.com/advanced-spectral-representations-audio-analysis/">AllYouNeedIsSound 3: Spectral Representations and Feature Extraction</a> appeared first on <a href="https://www.jhonatanlopez.com">Jhonatan López</a>.</p>
]]></description>
										<content:encoded><![CDATA[
<figure class="wp-block-image size-full"><img decoding="async" width="1024" height="1024" src="https://www.jhonatanlopez.com/wp-content/uploads/2025/03/allyouneedissound3.webp" alt="Three-dimensional waves representing spectral audio analysis in grey and light blue tones on a white background, with a stylized graphic equalizer." class="wp-image-4618" srcset="https://www.jhonatanlopez.com/wp-content/uploads/2025/03/allyouneedissound3.webp 1024w, https://www.jhonatanlopez.com/wp-content/uploads/2025/03/allyouneedissound3-300x300.webp 300w, https://www.jhonatanlopez.com/wp-content/uploads/2025/03/allyouneedissound3-100x100.webp 100w, https://www.jhonatanlopez.com/wp-content/uploads/2025/03/allyouneedissound3-600x600.webp 600w, https://www.jhonatanlopez.com/wp-content/uploads/2025/03/allyouneedissound3-150x150.webp 150w, https://www.jhonatanlopez.com/wp-content/uploads/2025/03/allyouneedissound3-768x768.webp 768w" sizes="(max-width: 1024px) 100vw, 1024px" /><figcaption class="wp-element-caption">A modern and minimalist interpretation of spectral audio analysis created with DALL·E.</figcaption></figure>



<p class="wp-block-paragraph">Have you ever wondered how machines understand the nuances of sound? In my previous post, we explored <a href="https://www.jhonatanlopez.com/spectral-analysis/">spectral analysis</a> and learned how spectrograms reveal the frequency content of audio signals using the Short-Time Fourier Transform (STFT). Now, let’s dive deeper into advanced spectral representations for audio analysis, including Mel Spectrograms, CQT, and HCQT, and show how they can be used for perceptual audio analysis and feature extraction. These tools are essential for building machine learning models for tasks like audio classification, a field I’m currently exploring.</p>



<h2 class="wp-block-heading">Why Feature Extraction?</h2>



<p class="wp-block-paragraph">Spectral analysis provides us a visual map of audio frequencies, but for machine learning, we need compact, meaningful features that capture the essence of sound. Raw spectrograms are rich but high-dimensional, making them inefficient for direct use in models. By refining them into perceptually relevant or musically meaningful representations, we can extract features that align with how we hear or interpret audio. This is crucial for applications like genre classification, pitch detection, or environmental sound recognition.</p>



<h2 class="wp-block-heading">Advanced Spectral Representations</h2>



<p class="wp-block-paragraph">Let’s explore three advanced spectral representations that address the limitations of STFT-based spectrograms: Mel Spectrograms, Constant-Q Transform (CQT), and Harmonic-CQT (HCQT). Each of these tools offers unique advantages for audio analysis and feature extraction.</p>



<h3 class="wp-block-heading">Mel Spectrogram (MEL) and Log-Mel Spectrogram (LMS)</h3>



<h4 class="wp-block-heading">What Are They?</h4>



<p class="wp-block-paragraph">The Mel Spectrogram adapts the STFT to the Mel scale, a perceptual scale of pitch that reflects how humans hear frequency differences (e.g., we’re more sensitive to changes at lower frequencies). It compresses the frequency axis into Mel bins, reducing dimensionality while prioritizing auditory perception. The Log-Mel Spectrogram takes this further by applying a logarithmic transformation to the amplitude, mimicking the logarithmic response of our ears to loudness.</p>



<h4 class="wp-block-heading">Why Use Them?</h4>



<ul class="wp-block-list">
<li><strong>Perceptual Relevance:</strong> Mel Spectrograms align with human hearing, making them ideal for speech and music analysis.</li>



<li><strong>Machine Learning Ready:</strong> Log-Mel Spectrograms are compact and widely used as input features for deep learning models.</li>
</ul>



<h4 class="wp-block-heading">Example in Python</h4>



<pre class="wp-block-code"><code>```python
import librosa
import librosa.display
import matplotlib.pyplot as plt
import numpy as np
from google.colab import drive

# Mount Google Drive
drive.mount('/content/drive')

# Load audio
y, sr = librosa.load('/content/drive/My Drive/audio_files/sample.wav')

# Compute Mel Spectrogram
S = librosa.feature.melspectrogram(y=y, sr=sr, n_mels=128)
S_db = librosa.power_to_db(S, ref=np.max)  # Log-Mel Spectrogram

# Plot
plt.figure(figsize=(14, 5))
librosa.display.specshow(S_db, sr=sr, x_axis='time', y_axis='mel')
plt.colorbar(format='%+2.0f dB')
plt.title('Violin Log-Mel Spectrogram')
plt.show()
```</code></pre>



<ul class="wp-block-list">
<li><strong>n_mels=128:</strong> Number of Mel bins (adjustable based on your needs).</li>



<li><strong>Output:</strong> Time vs. Mel frequency, with colour showing log-amplitude.</li>
</ul>



<figure class="wp-block-image size-large is-resized"><img decoding="async" width="1024" height="445" src="https://www.jhonatanlopez.com/wp-content/uploads/2025/02/violinLogMelSpectrogram-1024x445.png" alt="Log-Mel Spectrogram for advanced spectral representations in audio analysis" class="wp-image-4555" style="width:656px;height:auto" srcset="https://www.jhonatanlopez.com/wp-content/uploads/2025/02/violinLogMelSpectrogram-1024x445.png 1024w, https://www.jhonatanlopez.com/wp-content/uploads/2025/02/violinLogMelSpectrogram-600x261.png 600w, https://www.jhonatanlopez.com/wp-content/uploads/2025/02/violinLogMelSpectrogram-300x130.png 300w, https://www.jhonatanlopez.com/wp-content/uploads/2025/02/violinLogMelSpectrogram-768x334.png 768w, https://www.jhonatanlopez.com/wp-content/uploads/2025/02/violinLogMelSpectrogram.png 1081w" sizes="(max-width: 1024px) 100vw, 1024px" /><figcaption class="wp-element-caption"><em><strong>Figure 1: </strong>Example Log-Mel Spectrogram generated from an audio file using the code above. The x-axis represents time, and the y-axis shows frequency, giving a visual representation of the sound’s intensity over time.</em></figcaption></figure>



<h3 class="wp-block-heading">Constant-Q Transform (CQT)</h3>



<h4 class="wp-block-heading">What Is It?</h4>



<p class="wp-block-paragraph">The Constant-Q Transform (CQT) is an alternative to STFT that uses a logarithmic frequency scale, where the frequency resolution is constant relative to the center frequency (constant Q-factor). Unlike STFT’s fixed window size, CQT’s window size varies—longer for low frequencies, shorter for high ones.</p>



<h4 class="wp-block-heading">Why Use It?</h4>



<ul class="wp-block-list">
<li><strong>Musical Advantage:</strong> Its logarithmic scale matches the intervals of musical notes (e.g., octaves), making it perfect for pitch-related tasks like chord recognition or music transcription.</li>



<li><strong>Better Resolution:</strong> It captures low-frequency details (e.g., bass notes) better than STFT.</li>
</ul>



<h4 class="wp-block-heading">Example in Python</h4>



<p class="wp-block-paragraph">The following example was implemented following the code used before.</p>



<pre class="wp-block-code"><code>```python
# Compute CQT
C = librosa.cqt(y, sr=sr)
C_db = librosa.amplitude_to_db(abs(C), ref=np.max)

# Plot
plt.figure(figsize=(14, 5))
librosa.display.specshow(C_db, sr=sr, x_axis='time', y_axis='cqt_note')
plt.colorbar(format='%+2.0f dB')
plt.title('Violin Constant-Q Transform')
plt.show()
```</code></pre>



<p class="wp-block-paragraph"><strong>y_axis=&#8217;cqt_note&#8217;:</strong> Labels the y-axis with musical notes (e.g., C4, D4), emphasizing its musical focus.</p>



<figure class="wp-block-image size-large"><img loading="lazy" decoding="async" width="1024" height="452" src="https://www.jhonatanlopez.com/wp-content/uploads/2025/02/violinCQT-1024x452.png" alt="Constant-Q Transform (CQT) for advanced spectral representations in audio analysis." class="wp-image-4554" srcset="https://www.jhonatanlopez.com/wp-content/uploads/2025/02/violinCQT-1024x452.png 1024w, https://www.jhonatanlopez.com/wp-content/uploads/2025/02/violinCQT-600x265.png 600w, https://www.jhonatanlopez.com/wp-content/uploads/2025/02/violinCQT-300x133.png 300w, https://www.jhonatanlopez.com/wp-content/uploads/2025/02/violinCQT-768x339.png 768w, https://www.jhonatanlopez.com/wp-content/uploads/2025/02/violinCQT.png 1064w" sizes="auto, (max-width: 1024px) 100vw, 1024px" /><figcaption class="wp-element-caption"><em><strong>Figure 2:</strong> Example Constant-Q Transform generated from an audio file using the code above. The x-axis represents time, and the y-axis shows musical notation, giving a visual representation of the sound’s intensity over time.</em></figcaption></figure>



<h3 class="wp-block-heading">Harmonic-CQT (HCQT)</h3>



<h4 class="wp-block-heading">What Is It?</h4>



<p class="wp-block-paragraph">The Harmonic Constant-Q Transform (HCQT) extends CQT by analysing harmonic structures. It computes CQTs at multiple harmonic multiples (e.g., fundamental frequency and its overtones) and stacks them into a 3D representation.</p>



<h4 class="wp-block-heading">Why Use It?</h4>



<ul class="wp-block-list">
<li><strong>Pitch-Related Applications:</strong> HCQT excels at separating harmonic content (e.g., a piano’s notes) from noise or percussive elements, ideal for pitch detection or source separation.</li>



<li><strong>Research Edge:</strong> It’s advanced and less common, showcasing cutting-edge techniques.</li>
</ul>



<h4 class="wp-block-heading">Note on Implementation</h4>



<p class="wp-block-paragraph"><code>`Librosa` </code>doesn’t directly provide HCQT, but you can approximate it by computing CQTs for harmonic multiples manually or use external libraries like <code>`nnAudio`</code>. Here are simplified examples using both libraries:</p>



<p class="wp-block-paragraph">With <code>`Librosa`</code>:</p>



<pre class="wp-block-code"><code>```python
import librosa
import librosa.display
import matplotlib.pyplot as plt
import numpy as np

# Load audio file
y, sr = librosa.load('/content/drive/My Drive/audio_files/sample.wav', sr=22050)  # Replace with your file path
hop_length = 512  # Number of samples between successive frames
harmonics = &#91;1, 2, 3]  # Harmonics to analyze (fundamental + overtones)

# Compute HCQT for the fundamental (h=1)
fmin = librosa.note_to_hz('C1') * 1  # Convert note C1 to Hz (~32.7 Hz)
n_bins = 60  # Total bins (5 octaves: 60/12 = 5)

# Check Nyquist limit (prevents aliasing)
nyquist_limit = fmin * (2 ** (n_bins / 12)) 
if nyquist_limit &lt; sr / 2:
    # Compute Constant-Q Transform
    cqt = librosa.cqt(y, sr=sr, hop_length=hop_length, 
                     fmin=fmin, n_bins=n_bins, bins_per_octave=12)
else:
    raise ValueError("Nyquist limit exceeded! Adjust parameters.")

# Convert CQT magnitude to decibels (normalized to max amplitude)
cqt_db = librosa.amplitude_to_db(np.abs(cqt), ref=np.max)

# Generate CQT frequency axis (logarithmic scale)
frequencies = librosa.cqt_frequencies(n_bins=n_bins, fmin=fmin, bins_per_octave=12)

# Plot the spectrogram
plt.figure(figsize=(14, 5))
librosa.display.specshow(cqt_db, sr=sr, hop_length=hop_length,
                        y_axis='cqt_hz', x_axis='time',  # Log-frequency axis
                        fmin=fmin, bins_per_octave=12, 
                        vmin=-80, vmax=0)  # dB range and optional colourmap add , cmap='viridis'
plt.colorbar(format='%+2.0f dB', label='Amplitude (dB)')
plt.ylim(frequencies&#91;0], frequencies&#91;-1])  # Set frequency axis limits
plt.title('Violin Harmonic-CQT (Fundamental) - Librosa')
plt.xlabel('Time (s)')
plt.ylabel('Frequency (Hz)')
plt.show()
```</code></pre>



<details class="wp-block-details is-layout-flow wp-block-details-is-layout-flow"><summary><strong>Limitations:</strong></summary>
<p class="wp-block-paragraph">&nbsp; &#8211; Tedious manual setup.</p>



<p class="wp-block-paragraph">&nbsp; &#8211; No native harmonic stacking.</p>



<p class="wp-block-paragraph">&nbsp; &#8211; Limited to CPU computation.</p>
</details>



<p class="wp-block-paragraph"><em>For efficient HCQT computation, we use `nnAudio`, a PyTorch-based library that leverages GPU acceleration. First, install it:</em></p>



<pre class="wp-block-code"><code>```python
pip install nnAudio
```</code></pre>



<p class="wp-block-paragraph">Then, run the following code:</p>



<pre class="wp-block-code"><code>```python
import torch
from nnAudio.features.cqt import CQT
import matplotlib.pyplot as plt

# Parameters
sr = 22050  # Sample rate
hop_length = 512  # Hop size
n_bins = 60  # Number of frequency bins (reduced to avoid Nyquist issues)
fmin = 32.7  # Minimum frequency (C1 in Hz)
harmonics = &#91;1, 2, 3]  # Harmonics to compute

# Load audio (using librosa)
y, _ = librosa.load("/content/drive/My Drive/audio_files/sample.wav", sr=sr)

# Convert to PyTorch tensor
y_tensor = torch.tensor(y).float()

# Compute HCQT for each harmonic
hcqt = &#91;]
for h in harmonics:
    cqt = CQT(sr=sr, hop_length=hop_length, n_bins=n_bins,
              fmin=fmin * h, bins_per_octave=12, output_format='Magnitude')
    cqt_output = cqt(y_tensor)  # Shape: (1, n_bins, time)
    cqt_db = 20 * torch.log10(torch.clamp(cqt_output, min=1e-5))  # Avoid log(0)
    hcqt.append(cqt_db)

# Plot the fundamental harmonic
if hcqt:
    plt.figure(figsize=(14, 5))
    plt.imshow(hcqt&#91;0].squeeze().numpy(), aspect='auto', origin='lower', cmap='viridis', vmin=-80, vmax=0, interpolation='bilinear')
    plt.colorbar(format='%+2.0f dB')
    plt.title('Violin Harmonic-CQT (Fundamental) - nnAudio')
    plt.xlabel('Time')
    plt.ylabel('Frequency (bins)')
    plt.show()
```</code></pre>



<details class="wp-block-details is-layout-flow wp-block-details-is-layout-flow"><summary><strong>Advantages:</strong></summary>
<p class="wp-block-paragraph">&nbsp; &#8211; GPU Acceleration: Faster computation for large datasets.</p>



<p class="wp-block-paragraph">&nbsp; &#8211; Native Harmonic Support: Streamlined parameter setup.</p>



<p class="wp-block-paragraph">&nbsp; &#8211; PyTorch Integration: Direct compatibility with deep learning pipelines.</p>
</details>



<figure class="wp-block-image size-large"><img loading="lazy" decoding="async" width="1024" height="449" src="https://www.jhonatanlopez.com/wp-content/uploads/2025/02/harmonicCQTComparison1-1024x449.png" alt="Violin Harmonic-CQT (Fundamental) computed using Librosa, showing frequency and amplitude variations over time." class="wp-image-4550" srcset="https://www.jhonatanlopez.com/wp-content/uploads/2025/02/harmonicCQTComparison1-1024x449.png 1024w, https://www.jhonatanlopez.com/wp-content/uploads/2025/02/harmonicCQTComparison1-600x263.png 600w, https://www.jhonatanlopez.com/wp-content/uploads/2025/02/harmonicCQTComparison1-300x132.png 300w, https://www.jhonatanlopez.com/wp-content/uploads/2025/02/harmonicCQTComparison1-768x337.png 768w, https://www.jhonatanlopez.com/wp-content/uploads/2025/02/harmonicCQTComparison1.png 1072w" sizes="auto, (max-width: 1024px) 100vw, 1024px" /></figure>



<figure class="wp-block-image size-large"><img loading="lazy" decoding="async" width="1024" height="453" src="https://www.jhonatanlopez.com/wp-content/uploads/2025/02/harmonicCQTComparison2-1024x453.png" alt="Violin Harmonic-CQT (Fundamental) computed using nnAudio, showing frequency and amplitude variations over time." class="wp-image-4551" srcset="https://www.jhonatanlopez.com/wp-content/uploads/2025/02/harmonicCQTComparison2-1024x453.png 1024w, https://www.jhonatanlopez.com/wp-content/uploads/2025/02/harmonicCQTComparison2-600x265.png 600w, https://www.jhonatanlopez.com/wp-content/uploads/2025/02/harmonicCQTComparison2-300x133.png 300w, https://www.jhonatanlopez.com/wp-content/uploads/2025/02/harmonicCQTComparison2-768x340.png 768w, https://www.jhonatanlopez.com/wp-content/uploads/2025/02/harmonicCQTComparison2.png 1063w" sizes="auto, (max-width: 1024px) 100vw, 1024px" /><figcaption class="wp-element-caption"><strong>Figure 3: </strong>HCQT computed with librosa (top) vs. nnAudio (bottom). The nnAudio implementation offers cleaner harmonic separation due to GPU-optimized computation.<br><em>The axis are labelled different but basic programming configurations to plot are the same.</em></figcaption></figure>



<h2 class="wp-block-heading">What Do These Representations Tell Us?</h2>



<ul class="wp-block-list">
<li><strong>Mel/Log-Mel:</strong> Highlights perceptually significant frequencies (e.g., speech formants or musical timbre).</li>



<li><strong>CQT:</strong> Reveals musical structure (e.g., note transitions in a melody).</li>



<li><strong>HCQT:</strong> Isolates harmonic patterns (e.g., a chord’s overtones), distinguishing pitched sounds from noise.</li>
</ul>



<p class="wp-block-paragraph">These features are more targeted than raw STFT spectrograms, making them powerful inputs for machine learning models.</p>



<h2 class="wp-block-heading">Reflection</h2>



<p class="wp-block-paragraph">Exploring these spectral representations has been a transformative experience for me. Initially, I relied heavily on STFT, but discovering Mel Spectrograms showed me how aligning analysis with human perception could significantly boost classification accuracy—something I’m currently testing with various audio datasets. Implementing CQT was a revelation for its musical precision, though working with HCQT pushed my coding skills to the limit. I spent hours digging into research papers and experimenting with harmonic stacking to get it right. These challenges have deepened my understanding of audio feature extraction and increased my excitement for applying these techniques to machine learning models.</p>



<h2 class="wp-block-heading">Conclusion</h2>



<p class="wp-block-paragraph">Spectral representations like Mel Spectrograms, CQT, and HCQT take us beyond basic spectrograms, offering perceptually and musically relevant features for audio analysis.</p>



<p class="wp-block-paragraph">In this post, we’ve explored&nbsp;advanced spectral representations for audio analysis, including Mel Spectrograms, CQT, and HCQT, and seen how they can be used for audio analysis and feature extraction. These tools take us beyond waveforms and basic spectrograms, offering perceptually and musically relevant features that are essential for machine learning tasks.</p>



<h2 class="wp-block-heading">Additional Resources</h2>



<ul class="wp-block-list">
<li>Librosa Documentation:&nbsp;<a href="https://librosa.org/doc/">librosa.org/doc</a></li>



<li>nnAudio:&nbsp;<a href="https://kinwaicheuk.github.io/nnAudio/v0.2.0/index.html">nnAudio 0.2.0</a></li>



<li><a href="https://geoffroypeeters.github.io/deeplearning-101-audiomir_book">Deep Learning 101 for Audio-based MIR, ISMIR 2024 Tutorial</a>&nbsp;by Geoffroy Peeters et al. (2024).</li>



<li>Z. Rafii, &#8220;The Constant-Q Harmonic Coefficients: A timbre feature designed for music signals [Lecture Notes],&#8221; in IEEE Signal Processing Magazine, vol. 39, no. 3, pp. 90-96, May 2022, doi: 10.1109/MSP.2021.3138870. keywords: {Cepstral analysis;Instruments;Transforms;Speech recognition;Power system harmonics;Harmonic analysis;Feature extraction},</li>



<li>K. W. Cheuk, H. Anderson, K. Agres and D. Herremans, &#8220;nnAudio: An on-the-Fly GPU Audio to Spectrogram Conversion Toolbox Using 1D Convolutional Neural Networks,&#8221; in IEEE Access, vol. 8, pp. 161981-162003, 2020, doi: 10.1109/ACCESS.2020.3019084.</li>
</ul>
<p>The post <a href="https://www.jhonatanlopez.com/advanced-spectral-representations-audio-analysis/">AllYouNeedIsSound 3: Spectral Representations and Feature Extraction</a> appeared first on <a href="https://www.jhonatanlopez.com">Jhonatan López</a>.</p>
]]></content:encoded>
					
		
		
			</item>
	</channel>
</rss>
