6cd06a7a8607c54bddd1a96fc4bd4af9849d1bc9
OmniParser: Screen Parsing tool for Pure Vision Based GUI Agent
OmniParser is a comprehensive method for parsing user interface screenshots into structured and easy-to-understand elements, which significantly enhances the ability of GPT-4V to generate actions that can be accurately grounded in the corresponding regions of the interface.
Examples:
We put together a few simple examples in the demo.ipynb.
Gradio Demo
To run gradio demo, simply run:
python gradion_demo.py
📚 Citation
Our technical report can be found here. If you find our work useful, please consider citing our work:
@misc{lu2024omniparserpurevisionbased,
title={OmniParser for Pure Vision Based GUI Agent},
author={Yadong Lu and Jianwei Yang and Yelong Shen and Ahmed Awadallah},
year={2024},
eprint={2408.00203},
archivePrefix={arXiv},
primaryClass={cs.CV},
url={https://arxiv.org/abs/2408.00203},
}
Languages
Jupyter Notebook
50.1%
Python
36.8%
Shell
8.2%
PowerShell
4.5%
Dockerfile
0.2%
Other
0.2%
