语音转文字-Wio终端


文档摘要

语音转文字 - Wio终端 在本课程的这一部分,你将编写代码,使用语音服务将捕获的音频中的语音转换为文本。 将音频发送到语音服务 可以使用REST API将音频发送到语音服务。要使用语音服务,首先需要请求访问令牌,然后使用该令牌访问REST API。这些访问令牌每10分钟过期一次,因此你的代码应定期请求它们以确保它们始终是最新的。 任务 - 获取访问令牌 如果尚未打开,请打开项目 。 在 文件中添加以下库依赖项,以访问WiFi并处理JSON: 向 头文件中添加以下代码: 替换 and with the relevant values for your WiFi. Replace with the API key for your speech service resource.

语音转文字 - Wio终端

在本课程的这一部分,你将编写代码,使用语音服务将捕获的音频中的语音转换为文本。

将音频发送到语音服务

可以使用REST API将音频发送到语音服务。要使用语音服务,首先需要请求访问令牌,然后使用该令牌访问REST API。这些访问令牌每10分钟过期一次,因此你的代码应定期请求它们以确保它们始终是最新的。

任务 - 获取访问令牌

  1. 如果尚未打开,请打开项目 smart-timer

  2. platformio.ini 文件中添加以下库依赖项,以访问WiFi并处理JSON:

    seeed-studio/Seeed Arduino rpcWiFi @ 1.0.5 seeed-studio/Seeed Arduino rpcUnified @ 2.1.3 seeed-studio/Seeed_Arduino_mbedtls @ 3.0.1 seeed-studio/Seeed Arduino RTC @ 2.0.0 bblanchon/ArduinoJson @ 6.17.3
  3. config.h 头文件中添加以下代码:

    const char *SSID = "<SSID>"; const char *PASSWORD = "<PASSWORD>"; const char *SPEECH_API_KEY = "<API_KEY>"; const char *SPEECH_LOCATION = "<LOCATION>"; const char *LANGUAGE = "<LANGUAGE>"; const char *TOKEN_URL = "https://%s.api.cognitive.microsoft.com/sts/v1.0/issuetoken";

    替换 <SSID> and <PASSWORD> with the relevant values for your WiFi.

    Replace <API_KEY> with the API key for your speech service resource. Replace <LOCATION> with the location you used when you created the speech service resource.

    Replace <LANGUAGE> with the locale name for language you will be speaking in, for example en-GB for English, or zn-HK for Cantonese. You can find a list of the supported languages and their locale names in the Language and voice support documentation on Microsoft docs.

    The TOKEN_URL constant is the URL of the token issuer without the location. This will be combined with the location later to get the full URL.

  4. Just like connecting to Custom Vision, you will need to use an HTTPS connection to connect to the token issuing service. To the end of config.h,在 config.h 中添加以下代码:

    const char *TOKEN_CERTIFICATE = "-----BEGIN CERTIFICATE-----\r\n" "MIIF8zCCBNugAwIBAgIQAueRcfuAIek/4tmDg0xQwDANBgkqhkiG9w0BAQwFADBh\r\n" "MQswCQYDVQQGEwJVUzEVMBMGA1UEChMMRGlnaUNlcnQgSW5jMRkwFwYDVQQLExB3\r\n" "d3cuZGlnaWNlcnQuY29tMSAwHgYDVQQDExdEaWdpQ2VydCBHbG9iYWwgUm9vdCBH\r\n" "MjAeFw0yMDA3MjkxMjMwMDBaFw0yNDA2MjcyMzU5NTlaMFkxCzAJBgNVBAYTAlVT\r\n" "MR4wHAYDVQQKExVNaWNyb3NvZnQgQ29ycG9yYXRpb24xKjAoBgNVBAMTIU1pY3Jv\r\n" "c29mdCBBenVyZSBUTFMgSXNzdWluZyBDQSAwNjCCAiIwDQYJKoZIhvcNAQEBBQAD\r\n" "ggIPADCCAgoCggIBALVGARl56bx3KBUSGuPc4H5uoNFkFH4e7pvTCxRi4j/+z+Xb\r\n" "wjEz+5CipDOqjx9/jWjskL5dk7PaQkzItidsAAnDCW1leZBOIi68Lff1bjTeZgMY\r\n" "iwdRd3Y39b/lcGpiuP2d23W95YHkMMT8IlWosYIX0f4kYb62rphyfnAjYb/4Od99\r\n" "ThnhlAxGtfvSbXcBVIKCYfZgqRvV+5lReUnd1aNjRYVzPOoifgSx2fRyy1+pO1Uz\r\n" "aMMNnIOE71bVYW0A1hr19w7kOb0KkJXoALTDDj1ukUEDqQuBfBxReL5mXiu1O7WG\r\n" "0vltg0VZ/SZzctBsdBlx1BkmWYBW261KZgBivrql5ELTKKd8qgtHcLQA5fl6JB0Q\r\n" "gs5XDaWehN86Gps5JW8ArjGtjcWAIP+X8CQaWfaCnuRm6Bk/03PQWhgdi84qwA0s\r\n" "sRfFJwHUPTNSnE8EiGVk2frt0u8PG1pwSQsFuNJfcYIHEv1vOzP7uEOuDydsmCjh\r\n" "lxuoK2n5/2aVR3BMTu+p4+gl8alXoBycyLmj3J/PUgqD8SL5fTCUegGsdia/Sa60\r\n" "N2oV7vQ17wjMN+LXa2rjj/b4ZlZgXVojDmAjDwIRdDUujQu0RVsJqFLMzSIHpp2C\r\n" "Zp7mIoLrySay2YYBu7SiNwL95X6He2kS8eefBBHjzwW/9FxGqry57i71c2cDAgMB\r\n" "AAGjggGtMIIBqTAdBgNVHQ4EFgQU1cFnOsKjnfR3UltZEjgp5lVou6UwHwYDVR0j\r\n" "BBgwFoAUTiJUIBiV5uNu5g/6+rkS7QYXjzkwDgYDVR0PAQH/BAQDAgGGMB0GA1Ud\r\n" "JQQWMBQGCCsGAQUFBwMBBggrBgEFBQcDAjASBgNVHRMBAf8ECDAGAQH/AgEAMHYG\r\n" "CCsGAQUFBwEBBGowaDAkBggrBgEFBQcwAYYYaHR0cDovL29jc3AuZGlnaWNlcnQu\r\n" "Y29tMEAGCCsGAQUFBzAChjRodHRwOi8vY2FjZXJ0cy5kaWdpY2VydC5jb20vRGln\r\n" "aUNlcnRHbG9iYWxSb290RzIuY3J0MHsGA1UdHwR0MHIwN6A1oDOGMWh0dHA6Ly9j\r\n" "cmwzLmRpZ2ljZXJ0LmNvbS9EaWdpQ2VydEdsb2JhbFJvb3RHMi5jcmwwN6A1oDOG\r\n" "MWh0dHA6Ly9jcmw0LmRpZ2ljZXJ0LmNvbS9EaWdpQ2VydEdsb2JhbFJvb3RHMi5j\r\n" "cmwwHQYDVR0gBBYwFDAIBgZngQwBAgEwCAYGZ4EMAQICMBAGCSsGAQQBgjcVAQQD\r\n" "AgEAMA0GCSqGSIb3DQEBDAUAA4IBAQB2oWc93fB8esci/8esixj++N22meiGDjgF\r\n" "+rA2LUK5IOQOgcUSTGKSqF9lYfAxPjrqPjDCUPHCURv+26ad5P/BYtXtbmtxJWu+\r\n" "cS5BhMDPPeG3oPZwXRHBJFAkY4O4AF7RIAAUW6EzDflUoDHKv83zOiPfYGcpHc9s\r\n" "kxAInCedk7QSgXvMARjjOqdakor21DTmNIUotxo8kHv5hwRlGhBJwps6fEVi1Bt0\r\n" "trpM/3wYxlr473WSPUFZPgP1j519kLpWOJ8z09wxay+Br29irPcBYv0GMXlHqThy\r\n" "8y4m/HyTQeI2IMvMrQnwqPpY+rLIXyviI2vLoI+4xKE4Rn38ZZ8m\r\n" "-----END CERTIFICATE-----\r\n";

    这是连接到自定义视觉时使用的相同证书。

  5. main.cpp 文件顶部添加对WiFi头文件和config头文件的包含:

    #include <rpcWiFi.h> #include "config.h"
  6. main.cpp above the setup 函数中添加代码以连接到WiFi:

    void connectWiFi() { while (WiFi.status() != WL_CONNECTED) { Serial.println("Connecting to WiFi.."); WiFi.begin(SSID, PASSWORD); delay(500); } Serial.println("Connected!"); }
  7. 在建立串行连接后从 setup 函数中调用此函数:

    connectWiFi();
  8. src folder called speech_to_text.h 中创建一个新的头文件。在此头文件中,添加以下代码:

    #pragma once #include <Arduino.h> #include <ArduinoJson.h> #include <HTTPClient.h> #include <WiFiClientSecure.h> #include "config.h" #include "mic.h" class SpeechToText { public: private: }; SpeechToText speechToText;

    这包括了一些必要的头文件,用于HTTP连接、配置以及 mic.h header file, and defines a class called SpeechToText, before declaring an instance of that class that can be used later.

  9. Add the following 2 fields to the private 类的部分:

    WiFiClientSecure _token_client; String _access_token;

    _token_client is a WiFi Client that uses HTTPS and will be used to get the access token. This token will then be stored in _access_token.

  10. Add the following method to the private 部分:

    String getAccessToken() { char url[128]; sprintf(url, TOKEN_URL, SPEECH_LOCATION); HTTPClient httpClient; httpClient.begin(_token_client, url); httpClient.addHeader("Ocp-Apim-Subscription-Key", SPEECH_API_KEY); int httpResultCode = httpClient.POST("{}"); if (httpResultCode != 200) { Serial.println("Error getting access token, trying again..."); delay(10000); return getAccessToken(); } Serial.println("Got access token."); String result = httpClient.getString(); httpClient.end(); return result; }

    此代码使用语音资源的位置构建了令牌发行者的URL。然后创建了一个 HTTPClient to make the web request, setting it up to use the WiFi client configured with the token endpoints certificate. It sets the API key as a header for the call. It then makes a POST request to get the certificate, retrying if it gets any errors. Finally the access token is returned.

  11. To the public 部分,添加了一个获取访问令牌的方法。这将在以后的课程中用于将文本转换为语音。

    String AccessToken() { return _access_token; }
  12. public section, add an init 方法中添加设置令牌客户端的代码:

    void init() { _token_client.setCACert(TOKEN_CERTIFICATE); _access_token = getAccessToken(); }

    这设置了WiFi客户端上的证书,然后获取访问令牌。

  13. main.cpp 中,将这个新头文件添加到包含指令中:

    #include "speech_to_text.h"
  14. 初始化 SpeechToText class at the end of the setup function, after the mic.init call but before Ready 写入串行监视器:

    speechToText.init();

任务 - 从闪存读取音频

  1. 在本课程的早期部分,音频被记录到了闪存中。这段音频需要发送到语音服务的REST API,因此需要从闪存中读取。由于它太大而不能加载到内存缓冲区中,所以在 HTTPClient class that makes REST calls can stream data using an Arduino Stream - a class that can load data in small chunks, sending the chunks one at a time as part of the request. Every time you call read on a stream it returns the next block of data. An Arduino stream can be created that can read from the flash memory. Create a new file called flash_stream.h in the src 文件夹中添加以下代码:

    #pragma once #include <Arduino.h> #include <HTTPClient.h> #include <sfud.h> #include "config.h" class FlashStream : public Stream { public: virtual size_t write(uint8_t val) { } virtual int available() { } virtual int read() { } virtual int peek() { } private: };

    这声明了 FlashStream class, deriving from the Arduino Stream class. This is an abstract class - derived classes have to implement a few methods before the class can be instantiated, and these methods are defined in this class.

    ✅ Read more on Arduino Streams in the Arduino Stream documentation

  2. Add the following fields to the private 部分:

    size_t _pos; size_t _flash_address; const sfud_flash *_flash; byte _buffer[HTTP_TCP_BUFFER_SIZE];

    定义了一个临时缓冲区来存储从闪存中读取的数据,以及一些字段来存储从缓冲区读取时的当前位置、从闪存中读取的当前地址和闪存设备。

  3. private 部分中添加以下方法:

    void populateBuffer() { sfud_read(_flash, _flash_address, HTTP_TCP_BUFFER_SIZE, _buffer); _flash_address += HTTP_TCP_BUFFER_SIZE; _pos = 0; }

    此代码从闪存中当前地址读取数据并存储在缓冲区中。然后增加地址,以便下次调用时读取下一组内存。缓冲区的大小基于 HTTPClient will send to the REST API at one time.

    Erasing flash memory has to be done using the grain size, reading on the other hand does not.

  4. In the public 部分中的类,添加一个构造函数:

    FlashStream() { _pos = 0; _flash_address = 0; _flash = sfud_get_device_table() + 0; populateBuffer(); }

    此构造函数设置所有字段以从闪存块的开始位置读取,并加载第一个数据块到缓冲区。

  5. 实现 write 方法。此流只会读取数据,因此可以不做任何事情并返回0:

    virtual size_t write(uint8_t val) { return 0; }
  6. 实现 peek method. This returns the data at the current position without moving the stream along. Calling peek 方法。多次调用 peek 会一直返回相同的值,只要没有从流中读取数据。

    virtual int peek() { return _buffer[_pos]; }
  7. 实现 available 函数。此函数返回可以从流中读取的字节数,或如果流完成则返回-1。对于此类,最大可用字节数不会超过HTTPClient的块大小。当此流在HTTP客户端中使用时,它会调用此函数以查看有多少数据可用,然后请求发送到REST API的数据量。我们不希望每个块都超过HTTP客户端的块大小,所以如果有更多的数据可用,则返回块大小。如果少于块大小,则返回实际可用的数据量。一旦所有数据都被流传输完毕,返回-1。

    virtual int available() { int remaining = BUFFER_SIZE - ((_flash_address - HTTP_TCP_BUFFER_SIZE) + _pos); int bytes_available = min(HTTP_TCP_BUFFER_SIZE, remaining); if (bytes_available == 0) { bytes_available = -1; } return bytes_available; }
  8. 实现 read 方法以返回缓冲区中的下一个字节,同时增加位置。如果位置超过缓冲区的大小,则填充缓冲区并将位置重置为下一块从闪存中读取的数据。

    virtual int read() { int retVal = _buffer[_pos++]; if (_pos == HTTP_TCP_BUFFER_SIZE) { populateBuffer(); } return retVal; }
  9. speech_to_text.h 头文件中,为这个新头文件添加一个包含指令:

    #include "flash_stream.h"

任务 - 将语音转换为文本

  1. 通过将音频发送到语音服务的REST API,可以将语音转换为文本。此REST API具有与令牌发行者不同的证书,因此向 config.h 头文件添加以下代码以定义此证书:

    const char *SPEECH_CERTIFICATE = "-----BEGIN CERTIFICATE-----\r\n" "MIIF8zCCBNugAwIBAgIQCq+mxcpjxFFB6jvh98dTFzANBgkqhkiG9w0BAQwFADBh\r\n" "MQswCQYDVQQGEwJVUzEVMBMGA1UEChMMRGlnaUNlcnQgSW5jMRkwFwYDVQQLExB3\r\n" "d3cuZGlnaWNlcnQuY29tMSAwHgYDVQQDExdEaWdpQ2VydCBHbG9iYWwgUm9vdCBH\r\n" "MjAeFw0yMDA3MjkxMjMwMDBaFw0yNDA2MjcyMzU5NTlaMFkxCzAJBgNVBAYTAlVT\r\n" "MR4wHAYDVQQKExVNaWNyb3NvZnQgQ29ycG9yYXRpb24xKjAoBgNVBAMTIU1pY3Jv\r\n" "c29mdCBBenVyZSBUTFMgSXNzdWluZyBDQSAwMTCCAiIwDQYJKoZIhvcNAQEBBQAD\r\n" "ggIPADCCAgoCggIBAMedcDrkXufP7pxVm1FHLDNA9IjwHaMoaY8arqqZ4Gff4xyr\r\n" "RygnavXL7g12MPAx8Q6Dd9hfBzrfWxkF0Br2wIvlvkzW01naNVSkHp+OS3hL3W6n\r\n" "l/jYvZnVeJXjtsKYcXIf/6WtspcF5awlQ9LZJcjwaH7KoZuK+THpXCMtzD8XNVdm\r\n" "GW/JI0C/7U/E7evXn9XDio8SYkGSM63aLO5BtLCv092+1d4GGBSQYolRq+7Pd1kR\r\n" "EkWBPm0ywZ2Vb8GIS5DLrjelEkBnKCyy3B0yQud9dpVsiUeE7F5sY8Me96WVxQcb\r\n" "OyYdEY/j/9UpDlOG+vA+YgOvBhkKEjiqygVpP8EZoMMijephzg43b5Qi9r5UrvYo\r\n" "o19oR/8pf4HJNDPF0/FJwFVMW8PmCBLGstin3NE1+NeWTkGt0TzpHjgKyfaDP2tO\r\n" "4bCk1G7pP2kDFT7SYfc8xbgCkFQ2UCEXsaH/f5YmpLn4YPiNFCeeIida7xnfTvc4\r\n" "7IxyVccHHq1FzGygOqemrxEETKh8hvDR6eBdrBwmCHVgZrnAqnn93JtGyPLi6+cj\r\n" "WGVGtMZHwzVvX1HvSFG771sskcEjJxiQNQDQRWHEh3NxvNb7kFlAXnVdRkkvhjpR\r\n" "GchFhTAzqmwltdWhWDEyCMKC2x/mSZvZtlZGY+g37Y72qHzidwtyW7rBetZJAgMB\r\n" "AAGjggGtMIIBqTAdBgNVHQ4EFgQUDyBd16FXlduSzyvQx8J3BM5ygHYwHwYDVR0j\r\n" "BBgwFoAUTiJUIBiV5uNu5g/6+rkS7QYXjzkwDgYDVR0PAQH/BAQDAgGGMB0GA1Ud\r\n" "JQQWMBQGCCsGAQUFBwMBBggrBgEFBQcDAjASBgNVHRMBAf8ECDAGAQH/AgEAMHYG\r\n" "CCsGAQUFBwEBBGowaDAkBggrBgEFBQcwAYYYaHR0cDovL29jc3AuZGlnaWNlcnQu\r\n" "Y29tMEAGCCsGAQUFBzAChjRodHRwOi8vY2FjZXJ0cy5kaWdpY2VydC5jb20vRGln\r\n" "aUNlcnRHbG9iYWxSb290RzIuY3J0MHsGA1UdHwR0MHIwN6A1oDOGMWh0dHA6Ly9j\r\n" "cmwzLmRpZ2ljZXJ0LmNvbS9EaWdpQ2VydEdsb2JhbFJvb3RHMi5jcmwwN6A1oDOG\r\n" "MWh0dHA6Ly9jcmw0LmRpZ2ljZXJ0LmNvbS9EaWdpQ2VydEdsb2JhbFJvb3RHMi5j\r\n" "cmwwHQYDVR0gBBYwFDAIBgZngQwBAgEwCAYGZ4EMAQICMBAGCSsGAQQBgjcVAQQD\r\n" "AgEAMA0GCSqGSIb3DQEBDAUAA4IBAQAlFvNh7QgXVLAZSsNR2XRmIn9iS8OHFCBA\r\n" "WxKJoi8YYQafpMTkMqeuzoL3HWb1pYEipsDkhiMnrpfeYZEA7Lz7yqEEtfgHcEBs\r\n" "K9KcStQGGZRfmWU07hPXHnFz+5gTXqzCE2PBMlRgVUYJiA25mJPXfB00gDvGhtYa\r\n" "+mENwM9Bq1B9YYLyLjRtUz8cyGsdyTIG/bBM/Q9jcV8JGqMU/UjAdh1pFyTnnHEl\r\n" "Y59Npi7F87ZqYYJEHJM2LGD+le8VsHjgeWX2CJQko7klXvcizuZvUEDTjHaQcs2J\r\n" "+kPgfyMIOY1DMJ21NxOJ2xPRC/wAh/hzSBRVtoAnyuxtkZ4VjIOh\r\n" "-----END CERTIFICATE-----\r\n";
  2. 添加一个常量以定义没有位置的语音URL。稍后将结合位置和语言来获取完整的URL。

    const char *SPEECH_URL = "https://%s.stt.speech.microsoft.com/speech/recognition/conversation/cognitiveservices/v1?language=%s";
  3. speech_to_text.h header file, in the private section of the SpeechToText 类中,使用语音证书定义一个WiFi客户端字段:

    WiFiClientSecure _speech_client;
  4. init 方法中,设置此WiFi客户端上的证书:

    _speech_client.setCACert(SPEECH_CERTIFICATE);
  5. public section of the SpeechToText 类中添加以下代码,定义一个将语音转换为文本的方法:

    String convertSpeechToText() { }
  6. 向此方法中添加以下代码,创建一个使用配置了语音证书的WiFi客户端的HTTP客户端,并使用语音URL设置位置和语言:

    char url[128]; sprintf(url, SPEECH_URL, SPEECH_LOCATION, LANGUAGE); HTTPClient httpClient; httpClient.begin(_speech_client, url);
  7. 需要在连接上设置一些头部信息:

    httpClient.addHeader("Authorization", String("Bearer ") + _access_token); httpClient.addHeader("Content-Type", String("audio/wav; codecs=audio/pcm; samplerate=") + String(RATE)); httpClient.addHeader("Accept", "application/json;text/xml");

    设置使用访问令牌的授权头部、使用采样率的音频格式头部,并设置客户端期望结果为JSON。

  8. 之后,添加以下代码以进行REST API调用:

    Serial.println("Sending speech..."); FlashStream stream; int httpResponseCode = httpClient.sendRequest("POST", &stream, BUFFER_SIZE); Serial.println("Speech sent!");

    这创建了一个 FlashStream 并使用它将数据流式传输到REST API。

  9. 在此之下,添加以下代码:

    String text = ""; if (httpResponseCode == 200) { String result = httpClient.getString(); Serial.println(result); DynamicJsonDocument doc(1024); deserializeJson(doc, result.c_str()); JsonObject obj = doc.as<JsonObject>(); text = obj["DisplayText"].as<String>(); } else if (httpResponseCode == 401) { Serial.println("Access token expired, trying again with a new token"); _access_token = getAccessToken(); return convertSpeechToText(); } else { Serial.print("Failed to convert text to speech - error "); Serial.println(httpResponseCode); }

    这段代码检查响应码。

    如果是200,表示成功代码,然后检索结果,解码JSON,并将 DisplayText property is set into the text variable. This is the property that the text version of the speech is returned in.

    If the response code is 401, then the access token has expired (these tokens only last 10 minutes). A new access token is requested, and the call is made again.

    Otherwise, an error is sent to the serial monitor, and the text 留空。

  10. 在此方法末尾添加以下代码以关闭HTTP客户端并返回文本:

    httpClient.end(); return text;
  11. main.cpp call this new convertSpeechToText method in the processAudio 函数中,然后将语音输出到串行监视器:

    String text = speechToText.convertSpeechToText(); Serial.println(text);
  12. 构建此代码,将其上传到Wio终端并测试。一旦在串行监视器中看到 Ready,按下C按钮(最靠近电源开关的左侧按钮),然后说话。4秒的音频将被捕获并转换为文本。

    --- Available filters and text transformations: colorize, debug, default, direct, hexlify, log2file, nocontrol, printable, send_on_enter, time --- More details at http://bit.ly/pio-monitor-filters --- Miniterm on /dev/cu.usbmodem1101 9600,8,N,1 --- --- Quit: Ctrl+C | Menu: Ctrl+T | Help: Ctrl+T followed by Ctrl+H --- Connecting to WiFi.. Connected! Got access token. Ready. Starting recording... Finished recording Sending speech... Speech sent! {"RecognitionStatus":"Success","DisplayText":"Set a 2 minute and 27 second timer.","Offset":4700000,"Duration":35300000} Set a 2 minute and 27 second timer.

你可以在这份 code-speech-to-text/wio-terminal 文件夹中找到此代码。

你的语音转文字程序成功了!

声明:
本文件灏天文库团队进行了翻译。尽管我们力求准确,但请注意,翻译可能包含错误或不准确之处。原文档以其原始语言为准。我们不对因使用此翻译而产生的任何误解或误译负责。


作者与出处
原作者: microsoft
来源:microsoft
许可证:MIT
整理: 灏天文库整理
由灏天文库结构化整理,提供目录导航、全文检索与在线阅读,便于系统化学习
发布者: 作者: microsoft 转发
评论区 (0)
U