update readme install
This commit is contained in:
@@ -7,7 +7,7 @@
|
|||||||
[](https://arxiv.org/abs/2408.00203)
|
[](https://arxiv.org/abs/2408.00203)
|
||||||
[](https://opensource.org/licenses/MIT)
|
[](https://opensource.org/licenses/MIT)
|
||||||
|
|
||||||
📢 [[Project Page](https://microsoft.github.io/OmniParser/)] [[V2 Blog Post](https://www.microsoft.com/en-us/research/articles/omniparser-v2-turning-any-llm-into-a-computer-use-agent/)] [[Models V2](https://huggingface.co/microsoft/OmniParser-v2.0)] [[Models V1.5](https://huggingface.co/microsoft/OmniParser)] [[huggingface space (to be updated)](https://huggingface.co/spaces/microsoft/OmniParser)]
|
📢 [[Project Page](https://microsoft.github.io/OmniParser/)] [[V2 Blog Post](https://www.microsoft.com/en-us/research/articles/omniparser-v2-turning-any-llm-into-a-computer-use-agent/)] [[Models V2](https://huggingface.co/microsoft/OmniParser-v2.0)] [[Models V1.5](https://huggingface.co/microsoft/OmniParser)] [[HuggingFace Space Demo](https://huggingface.co/spaces/microsoft/OmniParser-v2)]
|
||||||
|
|
||||||
**OmniParser** is a comprehensive method for parsing user interface screenshots into structured and easy-to-understand elements, which significantly enhances the ability of GPT-4V to generate actions that can be accurately grounded in the corresponding regions of the interface.
|
**OmniParser** is a comprehensive method for parsing user interface screenshots into structured and easy-to-understand elements, which significantly enhances the ability of GPT-4V to generate actions that can be accurately grounded in the corresponding regions of the interface.
|
||||||
|
|
||||||
@@ -22,8 +22,9 @@
|
|||||||
- [2024/09] OmniParser achieves the best performance on [Windows Agent Arena](https://microsoft.github.io/WindowsAgentArena/)!
|
- [2024/09] OmniParser achieves the best performance on [Windows Agent Arena](https://microsoft.github.io/WindowsAgentArena/)!
|
||||||
|
|
||||||
## Install
|
## Install
|
||||||
Install environment:
|
First clone the repo, and then install environment:
|
||||||
```python
|
```python
|
||||||
|
cd OmniParser
|
||||||
conda create -n "omni" python==3.12
|
conda create -n "omni" python==3.12
|
||||||
conda activate omni
|
conda activate omni
|
||||||
pip install -r requirements.txt
|
pip install -r requirements.txt
|
||||||
@@ -31,10 +32,11 @@ pip install -r requirements.txt
|
|||||||
|
|
||||||
Ensure you have the V2 weights downloaded in weights folder (ensure caption weights folder is called icon_caption_florence). If not download them with:
|
Ensure you have the V2 weights downloaded in weights folder (ensure caption weights folder is called icon_caption_florence). If not download them with:
|
||||||
```
|
```
|
||||||
rm -rf weights/icon_detect weights/icon_caption weights/icon_caption_florence
|
# download the model checkpoints to local directory OmniParser/weights/
|
||||||
for f in icon_detect/{train_args.yaml,model.pt,model.yaml} icon_caption/{config.json,generation_config.json,model.safetensors}; do huggingface-cli download microsoft/OmniParser-v2.0 "$f" --local-dir weights; done
|
for f in icon_detect/{train_args.yaml,model.pt,model.yaml} icon_caption/{config.json,generation_config.json,model.safetensors}; do huggingface-cli download microsoft/OmniParser-v2.0 "$f" --local-dir weights; done
|
||||||
mv weights/icon_caption weights/icon_caption_florence
|
mv weights/icon_caption weights/icon_caption_florence
|
||||||
```
|
```
|
||||||
|
|
||||||
<!-- ## [deprecated]
|
<!-- ## [deprecated]
|
||||||
Then download the model ckpts files in: https://huggingface.co/microsoft/OmniParser, and put them under weights/, default folder structure is: weights/icon_detect, weights/icon_caption_florence, weights/icon_caption_blip2.
|
Then download the model ckpts files in: https://huggingface.co/microsoft/OmniParser, and put them under weights/, default folder structure is: weights/icon_detect, weights/icon_caption_florence, weights/icon_caption_blip2.
|
||||||
|
|
||||||
|
|||||||
Reference in New Issue
Block a user