microsoft-github-policy-service[bot] f4c34c92fc Microsoft mandatory file
2024-10-01 17:25:38 +00:00
2024-10-01 17:25:16 +00:00
2024-10-01 17:25:16 +00:00
2024-10-01 17:25:16 +00:00
2024-10-01 17:25:16 +00:00
2024-10-01 17:25:16 +00:00
2024-09-20 15:49:30 -07:00
2024-09-20 15:49:30 -07:00
2024-10-01 17:25:16 +00:00
2024-10-01 17:25:16 +00:00
2024-10-01 17:25:38 +00:00
2024-10-01 17:25:16 +00:00

OmniParser: Screen Parsing tool for Pure Vision Based GUI Agent

Logo arXiv License

OmniParser is a comprehensive method for parsing user interface screenshots into structured and easy-to-understand elements, which significantly enhances the ability of GPT-4V to generate actions that can be accurately grounded in the corresponding regions of the interface.

Install

conda create -n "omni" python==3.12
pip install -r requirements.txt

Examples:

We put together a few simple examples in the demo.ipynb.

Gradio Demo

To run gradio demo, simply run:

python gradion_demo.py

📚 Citation

Our technical report can be found here. If you find our work useful, please consider citing our work:

@misc{lu2024omniparserpurevisionbased,
      title={OmniParser for Pure Vision Based GUI Agent}, 
      author={Yadong Lu and Jianwei Yang and Yelong Shen and Ahmed Awadallah},
      year={2024},
      eprint={2408.00203},
      archivePrefix={arXiv},
      primaryClass={cs.CV},
      url={https://arxiv.org/abs/2408.00203}, 
}
Description
No description provided
Readme CC-BY-4.0 36 MiB
Languages
Jupyter Notebook 50.1%
Python 36.8%
Shell 8.2%
PowerShell 4.5%
Dockerfile 0.2%
Other 0.2%