← All work

CMU, 2024

Teaching an open model to read web pages

Multimodal models can read text and look at images, but they struggle with web pages: finding the right button, reading text inside a box, or guessing what a click will do. For our 11-777 (Multimodal Machine Learning) project, we tried to make a small open model, LLaVA-1.5-7B, better at these web tasks, and measured it on the VisualWebBench benchmark. This was fall 2024, when open models were far weaker at this than they are today.

With Akshay Badagabettu, Aayush Shah and Sai Yarlagadda, equal contributors. Multimodal Machine Learning (11-777), Carnegie Mellon, fall 2024.

Action prediction, before
10.0%
Action prediction, after
77.9%
Final report · 23 pages Download PDF

The problem

VisualWebBench is 1,500 webpage screenshots from 139 real websites, with seven tasks: captioning a page, answering questions about it, reading the main heading, reading text inside a marked box (Element OCR), picking the right element from a description, choosing what to click for an instruction, and predicting which page a click leads to (Action Prediction).

Out of the box, LLaVA-1.5-7B averaged 13.07 across the seven tasks. Looking at its mistakes, it tended to focus on big, bold text in the middle of the page instead of the box it was asked about, and it often ignored the requested answer format.

What we did

First we ran baselines on several models (LLaVA, TroL, Phantom, InternLM and BLIP-2), including text-only and image-only versions, to see what each modality contributes.

Version 1 was fine-tuning. We trained LLaVA on a random subset of MultiUI, a large dataset of web-UI instructions, using LoRA (a cheap way to fine-tune by training small add-on matrices instead of the whole model): 12.5k samples for 5 epochs, then 2 more epochs with another 60k.

Version 2, which I led as an equal contributor, changed nothing in the weights. It changed what we send the model: task-specific prompts with examples and clear output rules (for example, answer options listed as A, B, C), and bounding-box tricks, first cropping the image to the marked box, then appending the box below the full page so the model sees both.

Diagram: LLaVA 1.5 fine-tuned on the MultiUI dataset becomes Version 1; prompt enhancements turn it into Version 2; both are evaluated on VisualWebBench.
Our approach. Version 1 fine-tunes LLaVA on MultiUI. Version 2 adds sample outputs to the prompts and crops to the region of interest.

Results

Fine-tuning alone moved the average from 13.07 to 14.62. Adding the Version 2 prompts and preprocessing took it to 27.38.

The big jumps were in Action Prediction (9.96 to 77.94) and Element OCR (6.32 to 54.82). 77.94 was the highest Action Prediction score in our comparison table, above the closed models, including Gemini 1.5 Pro (74.4) and GPT-4 Vision (67.6).

Bar chart of LLaVA scores on seven VisualWebBench tasks for four versions: no training, checkpoint 1, full tuning, and full tuning plus prompt enhancements.
Scores per task as the model changed. The orange bars (fine-tuning plus prompts) jump in Element OCR and Action Prediction.
Table comparing closed-source and open-source models on the seven VisualWebBench tasks, including our LLaVA Version 1 and Version 2 rows.
Full comparison table. Our Version 2 row reaches 77.94 on Action Prediction and 54.82 on Element OCR.

What I took away

Cropping to the box helped in clear-cut cases, but it also raised a fair question: were we turning a multimodal task into a plain OCR task? Appending the box below the full page kept both the local and the page-wide context, and that fit the spirit of the benchmark better.

The gains were uneven. WebQA and captioning stayed well behind the best models, and hand-written prompts do not scale, so the report suggests automating prompt search as next steps.

Attention maps for each generated token next to the A-Z Animals webpage screenshot the model was asked about.
One hard question the fine-tuned model got right: which extra platform the site mentions (answer: YouTube). Left: attention maps per output token.

Tools LLaVA-1.5-7B, LoRA, MultiUI, VisualWebBench

Report

Loading…