# Audio link unique identifier

**URL:** <https://community.wanikani.com/t/audio-link-unique-identifier/8998>\
**Category:** API And Third-Party Apps\
**Created:** [July 5, 2015, 6:06am UTC](https://community.wanikani.com/t/audio-link-unique-identifier/8998 "2015-07-05T06:06:37Z")\
**Posts on this page:** 12\
**Page:** 1

<div class="post-metadata">

**Author:** ![SleepySpider](https://sea1.discourse-cdn.com/wanikanicommunity/user_avatar/community.wanikani.com/sleepyspider/32/4660_2.png) [@SleepySpider](https://community.wanikani.com/u/SleepySpider)\
**Post date:** [July 4, 2015, 10:06pm UTC](https://community.wanikani.com/t/audio-link-unique-identifier/8998/1 "2015-07-04T22:06:37Z")

</div>

I need to somehow grab the url for a particular vocab word to play the audio.  
If I look at the page source it’s something like this:&nbsp;[https://s3.amazonaws.com/s3.wanikani.com/audio/cf1d465792f380a6a2841c0732545454be2b3041.mp3](https://s3.amazonaws.com/s3.wanikani.com/audio/cf1d465792f380a6a2841c0732545454be2b3041.mp3)" type=“audio/mpeg”\>  
  
Which is fine, except I don’t think the API returns that unique identifier when you request the vocab data.  
So, the question is: How do I get this unique identifier? Or, if I have a vocab word, how to I get the correct url to play the mp3?  
  
I tried to see if maybe I could use Jisho’s audio, but they also use unique identifiers, of which I don’t know how to map.  
  
As of right now, I sorta have [japanesepod101.com](http://japanesepod101.com)’s audio assets working, but I have to pass in both a kanji and a kana reading to get the audio. Which I don’t understand why.  
  
Like this:  
[http://assets.languagepod101.com/dictionary/japanese/audiomp3.php?kanji=大した&kana=たいし](http://assets.languagepod101.com/dictionary/japanese/audiomp3.php?kanji=%E5%A4%A7%E3%81%97%E3%81%9F&kana=%E3%81%9F%E3%81%84%E3%81%97)た  
  
You can also use an id, but those are all I know. I don’t know if there are any other ways, and I don’t know how to figure out what the interface is to audiomp3.php to figure out if there are better ways to query the audio.  
  
Anybody have any ideas or suggestions?

---

<div class="post-metadata">

**Author:** ![gizmo](https://cdck-file-uploads-global.s3.dualstack.us-west-2.amazonaws.com/wanikanicommunity/original/3X/f/d/fd4c154120954695f788402f3bcf4e616499bc2d.png) [@gizmo](https://community.wanikani.com/u/gizmo)\
**Post date:** [July 4, 2015, 10:31pm UTC](https://community.wanikani.com/t/audio-link-unique-identifier/8998/2 "2015-07-04T22:31:43Z")

</div>

You could scrape the WaniKani vocab pages for the URL. The vocab pages themselves are deterministic at&nbsp;[https://www.wanikani.com/vocabulary/\<](https://www.wanikani.com/vocabulary/<)kanji\>. From looking at the source, you could then simply parse for any “source” tags and take their “src” attribute, which should give you one link for a .ogg, and one link for a .mp3  
  
As this is not offered through the API, you might want to check with the devolpers that this is something they are okay with you doing. Or at least rate-limit yourself.

---

<div class="post-metadata">

**Author:** ![zosiu](https://sea1.discourse-cdn.com/wanikanicommunity/user_avatar/community.wanikani.com/zosiu/32/224462_2.png) [@zosiu](https://community.wanikani.com/u/zosiu)\
**Post date:** [July 5, 2015, 5:36am UTC](https://community.wanikani.com/t/audio-link-unique-identifier/8998/3 "2015-07-05T05:36:47Z")

</div>

You could also use google TTS:&nbsp;  
`format: http://translate.google.com/translate_tts?ie=UTF-8&amp;q=YOUR_WORD&amp;tl=jaexample: http://translate.google.com/translate_tts?ie=UTF-8&q=調子はどうですか&tl=ja `

---

<div class="post-metadata">

**Author:** ![anon91083167](https://cdck-file-uploads-global.s3.dualstack.us-west-2.amazonaws.com/wanikanicommunity/original/3X/f/d/fd4c154120954695f788402f3bcf4e616499bc2d.png) [@anon91083167](https://community.wanikani.com/u/anon91083167)\
**Post date:** [July 5, 2015, 6:39am UTC](https://community.wanikani.com/t/audio-link-unique-identifier/8998/4 "2015-07-05T06:39:14Z")

</div>

Last time we talked about this, the thread got blocked ^^  
[/t/Humble-Request-Wanikani-to-Anki-with-Audio-Files/7314/1](https://community.wanikani.com/t/Humble-Request-Wanikani-to-Anki-with-Audio-Files/7314/1)

---

<div class="post-metadata">

**Author:** ![xMunch](https://sea1.discourse-cdn.com/wanikanicommunity/user_avatar/community.wanikani.com/xmunch/32/4389_2.png) [@xMunch](https://community.wanikani.com/u/xMunch)\
**Post date:** [July 5, 2015, 7:11am UTC](https://community.wanikani.com/t/audio-link-unique-identifier/8998/5 "2015-07-05T07:11:08Z")

</div>

> MarioRash said... Last time we talked about this, the thread got blocked ^^  
> [/t/Humble-Request-Wanikani-to-Anki-with-Audio-Files/7314/1](https://community.wanikani.com/t/Humble-Request-Wanikani-to-Anki-with-Audio-Files/7314/1)

&nbsp;They didn't say no ;^).  
  
  
  
I could throw a crawler together in the morning after I wake up&nbsp;if you'd like.  
  
If you want to do it yourself, I would suggest getting links to the individual vocab page through here:[https://www.wanikani.com/lattice/vocabulary/status](https://www.wanikani.com/lattice/vocabulary/status)&nbsp;then as said above parse audio url. &nbsp;Easy.

---

<div class="post-metadata">

**Author:** ![SleepySpider](https://sea1.discourse-cdn.com/wanikanicommunity/user_avatar/community.wanikani.com/sleepyspider/32/4660_2.png) [@SleepySpider](https://community.wanikani.com/u/SleepySpider)\
**Post date:** [July 5, 2015, 3:16pm UTC](https://community.wanikani.com/t/audio-link-unique-identifier/8998/6 "2015-07-05T15:16:42Z")

</div>

I don’t actually&nbsp;want the audio, just the links to the audio.  
  
But…&nbsp;I don’t want to do a crawler because every time new vocab words are added, I’d have to recrawl.  
I sorta want it generalizable. In other words, I want to be able to use audio from wanikani, and if wanikani doesn’t have it, then I go elsewhere. Since it’s sorta a wanikani chrome extension.  
  
Or, if I can’t use wanikani at all, just use a generic source.  
The google translate seems like an excellent option to start with. Thanks! If it turns out that I don’t like it, maybe I’ll figure something else out.  
  
It looks like, from that other mentioned thread, that wanikani left the audio stuff off the api intentionally. And I don’t want to rape wanikani’s frontend (or&nbsp;backend)&nbsp;to get what I want, it would make me feel bad. ;-p

---

<div class="post-metadata">

**Author:** ![xMunch](https://sea1.discourse-cdn.com/wanikanicommunity/user_avatar/community.wanikani.com/xmunch/32/4389_2.png) [@xMunch](https://community.wanikani.com/u/xMunch)\
**Post date:** [July 5, 2015, 3:35pm UTC](https://community.wanikani.com/t/audio-link-unique-identifier/8998/7 "2015-07-05T15:35:24Z")

</div>

Oh… &nbsp;I guess I’ll stop my crawling. &nbsp;I was working at a pretty conservative rate, ~40&nbsp;requests per minute.  
  
Edit: &nbsp;Just curious, what are you trying to do with this.

---

<div class="post-metadata">

**Author:** ![SleepySpider](https://sea1.discourse-cdn.com/wanikanicommunity/user_avatar/community.wanikani.com/sleepyspider/32/4660_2.png) [@SleepySpider](https://community.wanikani.com/u/SleepySpider)\
**Post date:** [July 5, 2015, 4:14pm UTC](https://community.wanikani.com/t/audio-link-unique-identifier/8998/8 "2015-07-05T16:14:35Z")

</div>

Thanks for the crawling effort!  
  
Well, I was going to suprise everyone, but I guess I’ll tell you guys here since you guys helped me out.  
I’m adding quite a few&nbsp;features to the&nbsp;wanikanify extension. Here’s the current changelist I’m working on.  
&nbsp;I’m done with most of them except some error checking and chrome sync.  
  
If you’re not sure what wanikanify is, it basically replaces english words with japanese vocab words from wanikani on the fly as you browse the internets. Clicking the word changes it back and forth from english/japanese. This occurs depending upon what level you’re currently at.

Changelist:

-Clicking on a translated vocab word will cause audio to play using Google TTS. (Which now works btw, but it’s not perfect as it can play the wrong pronunciation).

-User can now import large amounts of vocab words using google spreadsheets to supplement wanikani vocab. (Also works, but still needs some more testing)  

-User can now override wanikani vocab entries AND “Google spreadsheets import” using the “custom vocab box”. (Also works) For example, “time” gets translated as “〜回”, which is silly. So you can use this to override it to just “回” if you want.  

-Settings for wanikanify now are persistence across computers if user has Chrome’s sync functionality enabled. (In progress)  
  
I’m not the original author of wanikanify, but I forked his repo and when I’m done he’ll merge my changes back in, (he said my changes sounded cool)&nbsp;then we’ll release the changes to the google chrome extension store thingy.  
  
An interesting side effect of the spreadsheet feature, wanikanify can also&nbsp;be used with other languages as well, not just japanese. I’ll have to disable the requirement of adding an api key though, but there’s no reason it can’t translate english words to german, for example.

---

<div class="post-metadata">

**Author:** ![xMunch](https://sea1.discourse-cdn.com/wanikanicommunity/user_avatar/community.wanikani.com/xmunch/32/4389_2.png) [@xMunch](https://community.wanikani.com/u/xMunch)\
**Post date:** [July 5, 2015, 4:44pm UTC](https://community.wanikani.com/t/audio-link-unique-identifier/8998/9 "2015-07-05T16:44:26Z")

</div>

Just saying, there is a way you could only crawl new vocab. &nbsp;It would be similar to how I store UIDs on my userscript. &nbsp;You find which uids you don’t have an work from there. &nbsp;I could throw something together for you if you’d like.

---

<div class="post-metadata">

**Author:** ![SleepySpider](https://sea1.discourse-cdn.com/wanikanicommunity/user_avatar/community.wanikani.com/sleepyspider/32/4660_2.png) [@SleepySpider](https://community.wanikani.com/u/SleepySpider)\
**Post date:** [July 5, 2015, 6:50pm UTC](https://community.wanikani.com/t/audio-link-unique-identifier/8998/10 "2015-07-05T18:50:50Z")

</div>

Yea sure, I think that would be very&nbsp;helpful. thanks!  
I might end up using a combination of these or something. I’m not quite sure yet.  
I’m not really a web developer, so getting help with this would be good, if it’s not too much work.  
  
Thanks!

---

<div class="post-metadata">

**Author:** ![xMunch](https://sea1.discourse-cdn.com/wanikanicommunity/user_avatar/community.wanikani.com/xmunch/32/4389_2.png) [@xMunch](https://community.wanikani.com/u/xMunch)\
**Post date:** [July 5, 2015, 7:11pm UTC](https://community.wanikani.com/t/audio-link-unique-identifier/8998/11 "2015-07-05T19:11:51Z")

</div>

> aragonsr said... Yea sure, I think that would be very&nbsp;helpful. thanks!  
> I might end up using a combination of these or something. I'm not quite sure yet.  
> I'm not really a web developer, so getting help with this would be good, if it's not too much work.  
>   
> Thanks!

[https://gist.github.com/xMunch/3dd4221cead8c9572faf](https://gist.github.com/xMunch/3dd4221cead8c9572faf)&nbsp;-- JS audio file(It's only an object with the audio links; vocab:link format)  
[https://gist.github.com/xMunch/26a79c44a66d0083f00d](https://gist.github.com/xMunch/26a79c44a66d0083f00d)&nbsp;-- Crawler to update links.

---

<div class="post-metadata">

**Author:** ![SleepySpider](https://sea1.discourse-cdn.com/wanikanicommunity/user_avatar/community.wanikani.com/sleepyspider/32/4660_2.png) [@SleepySpider](https://community.wanikani.com/u/SleepySpider)\
**Post date:** [July 6, 2015, 2:10am UTC](https://community.wanikani.com/t/audio-link-unique-identifier/8998/12 "2015-07-06T02:10:53Z")

</div>

Great thanks!
