I think it's great to use JavaScript to control and combine the best of both worlds: accessibility APIs plus screen scraping / selective screencasting / pattern recognition / computer vision.
For example, you could use the accessibility APIs to find the screen position of the video window in the Skype application, perform facial recognition and tracking, and screencast the video onto a texture of a VR chat application.
For example, you could use the accessibility APIs to find the screen position of the video window in the Skype application, perform facial recognition and tracking, and screencast the video onto a texture of a VR chat application.