I don't think you can know the data consumption per say but its likely very very low. It sends a low bit rate version of the text you speak/dictate to the servers for processing and then back. We are probably talking a few K. Thats how nuance does it and this is their technology.