Decoding and Listening to Skype Voicemail .dat Files

October 2022 (1/3) | Listening to Skype Voicemail .dat files
April 2024 (2/3) | Still trying to listen to Skype Voicemails…

TL;DR, Scott Nickell discovered the final piece of the puzzle and gave me a heads up on it. See repository with updated README.md.

The end of the trilogy has just began…

4 Years Later…

I have this habit of Subconscious Processing, or “Letting It Simmer”, where when I have a problem that has no immediate solution I trust that your mind will piece together the answer over time. I’ve done this with many projects over time, and when the light bulb hits I make a couple notes here or there and maybe decide to pick the problem back up again. I’ve been doing this with the Skype Voicemail problem for 4 years now.

Recently I’ve been working tiring 12-hour shifts at an Amazon Data Center, making sure break-fix tickets get properly triaged, resolved, or moved to their relevant teams. The hours are cumbersome, the thinking is immense, the patience for diagnostics to finish is immeasurable. It’s a thankless corporate job, and every morning I depart before sunrise to speed through a drive-through, get myself a bagel, and get home to check messages and get to sleep.

Sitting in a special folder in my inbox is a message from Scott Nickell.Sitting in a special folder in my inbox is a message from Scott Nickell.

Hi,

I’m in a similar boat: have a bunch of .dat files that I kept thinking I’d be able to play them back easily, yet not.

I’ve leveraged a work malware reverse engineering tool that’s AI driven against multiple Skype binaries and the voicemail files to figure out the format and playback methodology to no avail. I was able to get it to dump out a handful of options it used (g729, iLBC, PCM, and so on) but nothing seemed to work with FFMPEG.

(That said, one interesting nugget I was able to get out of this was that the filename is constructed in a specific way, Unix timestamp being one part like you described, but the random number is in part random but it has bit more going on, though, was never able to figure out why it was so random in the first place. If I hear back from you, I’ll see if I can dig it up.)

WHY I’m reaching out: I got frustrated and just threw 5 of my .dat files against Anthropic’s Claude (Sonnet 5.5) and in like a minute it coughed up .wav versions and explained what format it is. Output below:

“These were recorded VoIP calls, and I converted all five to standard WAV files that play in any player.

What the files are: each .dat is a stream of length-prefixed RTP packets carrying G.729 audio (8 kHz, 20 ms per packet). No packets were missing in any file. Ordinary players can’t open that format directly, which is why they didn’t play as-is.

How I converted them: I stripped the framing and RTP headers, then decoded the G.729 payload with ffmpeg. To repeat this on other files from the same source, the command after extracting the payload is ffmpeg -f g729 -i payload.g729 out.wav.”

I further pressed it for how to get the g729 payload out and it spat out a python3 script and what it did:

“Split the records. Each record starts with a 2-byte big-endian length (00 20, which is 32 bytes). Records with length 0 are padding and get skipped.

Strip the RTP header. Each record is an RTP packet. Its first byte 0x80 means version 2, and the second byte 0x12 (18) is the payload type for G.729. The header is 12 bytes, plus 4 bytes per CSRC entry if any, which these files don’t have.

Join the payloads. The remaining 20 bytes of each packet are two 10-byte G.729 frames. Concatenating them across all packets gives a raw .g729 file.”

For the python script: https://claude.ai/public/artifacts/753500d3-2561-46c3-b4e9-bd0e7ae219d4

So, need python3 and ffmpeg installed and that script will shoot out both the g729 and the final wav.

Please feel free to reach out, warmest regards, and thank you for your research and sharing.

Scott

Doubtful it would work, as I swear I ran that bare data thorough ffmpeg before, and even tried vibe-coding this solution with my near-unlimited Token Allotment to failure months ago. Regardless, I verified I had ffmpeg in my path, gave the python script a try, and watched ffmpeg not choke on the output. Upon listening to the finalized wav files I breathed a sigh of relief.

The road less traveled…

I realize later that I was on the right path myself, I just didn’t put the pieces together.

My first approach was using ffmpeg and all supported demuxers to see if I could force this out. I even started the dive into WebRTC, but that wasn’t deep enough. If I descended down this rabbit hole I would’ve found more supported codecs aside from 

My second approach revolved around using Audio Formats that Skype authors would’ve used, and finally postulated running Skype with a disassembly program to step through and see exactly what the program would be doing.

Even breaking down the provided write-up makes perfect sense:

00 20 80 12 00 ee 00 00 94 c0 00 00 00 00 bd e3 d4 a1 d9 27 f0 81 95 0a 29 28 54 f4 a0 17 10 0e e9 ef

00 20 = length of packet
80 12 = RTP with g.729 encoding
00 ee = Sequence number
00 00 94 c0 = Timestamp
00 00 00 00 = SSRC Identifier
bd e3 d4 a1 d9 27 f0 81 95 0a 29 28 54 f4 a0 17 10 0e e9 ef = Payload

It’s RTP with a proprietary Skype length header. g.729 was misapplied to the data presented to when I ripped all the data out and presented it to ffmpeg 4 years ago.

Regardless, everything else makes sense, and I’m very thankful for it.

Epilogue

To me now this main project is completed. I’ve posted the Python script and instructions to the github repo.

My huge thanks go out to Scott for finding not only the curiosity and drive, but following that drive all the way to the solution, and finally with passing along the knowledge.

Leave a Reply