iFANN
    Buscar en iFANN...
    Iniciar sesión
    Inicio
    Noticias
    Vídeos
    Fotos
    GIFs
    Explorar
    Encuestas
    Premios
    iFAMOUS
    Wiki
    Anime
    Salas
    Notificaciones
    Mensajes
    Guardados
    Perfil
    WikiPremiosiFAMOUSClasificacionesSectoresRecompensas para creadoresRecompensas para usuariosTérminosPrivacidadNormas de la comunidadRetirada / DMCAAyudaDesarrolladores

    © 2026 iFANN

    Inicio
    Buscar
    Mensajes
    Alertas
    Perfil
    Foto
    Nate
    Nate@nate_5122mo
    💭Tech💭AI
    PixelRAG web screenshots Beat Text UC Berkeley

    @nate_512A new approach to web scraping for RAG systems: researchers at UC Berkeley have open-sourced PixelRAG, a tool that bypasses HTML parsing entirely. Instead of extracting text from a page and embedding chunks, it captures full-page screenshots and uses visual search over millions of rendered pages. The GitHub repository lists authors Yichuan Wang, Zhifei Li, Zirui Wang, Paul Teleltche, Lesheng Jin, Matei Zaharia, Joseph E. Gonzalez, and Sewon Min. The tool can be installed via pip and includes code for rendering pages to

    Ver publicación original

    PixelRAG web screenshots Beat Text UC Berkeley

    Foto de @nate_512· Jun 20, 2026· Tech

    Sobre esta foto

    A black cat mascot holding a screenshot icon sits beside the bold PixelRAG logo with a blue pixel accent. The tagline reads “Web Screenshots Beat Text” and the subhead explains “Visual search at scale over millions of rendered pages.” Below, the GitHub repo header lists authors Yichuan Wang, Zhifei Li, Zirui Wang, Paul Teleltche, Lesheng Jin, Matei Zaharia, Joseph E. Gonzalez, and Sewon Min, plus CI, demo, and live status badges. A terminal-style box shows the pip install command, followed by code snippets for rendering a page to screenshots and searching a visual index of 8.28M Wikipedia pages.

    Ver todas las fotos de Tech

    ?

    Aún no hay comentarios. ¡Sé el primero!

    Más fotos de Tech

    Ver todas las fotos de Tech
    Elon Musk AGI warning vs nuclear weaponsElon Musk AGI warning vs nuclear weaponsIndian workers record jobs for robot trainingIndian workers record jobs for robot trainingVenus Aerospace Rotating Detonation Engine Test StandVenus Aerospace Rotating Detonation Engine Test StandiPhone 18 Pro and iPhone 18 Pro Max pricingiPhone 18 Pro and iPhone 18 Pro Max pricingApple Samsung crease memeApple Samsung crease memeSwitzerland switches to open-source2Switzerland switches to open-sourceApple iPhone gaming controllerApple iPhone gaming controllerNew Designers 2026New Designers 2026unique soccer team momentunique soccer team momentGoogle Workspace voice featuresGoogle Workspace voice featuresElon Musk AI warningElon Musk AI warningTesla next-generation RoadsterTesla next-generation RoadsterOriginal content program application reviewOriginal content program application reviewSamsung Galaxy Z Fold8 Tim Cook ad2Samsung Galaxy Z Fold8 Tim Cook adiPhone Duo pricing tiers2iPhone Duo pricing tiersNikola Tesla solitude quoteNikola Tesla solitude quote
    Foto
    Nate
    Nate@nate_5122mo
    💭Tech💭AI
    PixelRAG web screenshots Beat Text UC Berkeley

    @nate_512A new approach to web scraping for RAG systems: researchers at UC Berkeley have open-sourced PixelRAG, a tool that bypasses HTML parsing entirely. Instead of extracting text from a page and embedding chunks, it captures full-page screenshots and uses visual search over millions of rendered pages. The GitHub repository lists authors Yichuan Wang, Zhifei Li, Zirui Wang, Paul Teleltche, Lesheng Jin, Matei Zaharia, Joseph E. Gonzalez, and Sewon Min. The tool can be installed via pip and includes code for rendering pages to

    Ver publicación original

    PixelRAG web screenshots Beat Text UC Berkeley

    Foto de @nate_512· Jun 20, 2026· Tech

    Sobre esta foto

    A black cat mascot holding a screenshot icon sits beside the bold PixelRAG logo with a blue pixel accent. The tagline reads “Web Screenshots Beat Text” and the subhead explains “Visual search at scale over millions of rendered pages.” Below, the GitHub repo header lists authors Yichuan Wang, Zhifei Li, Zirui Wang, Paul Teleltche, Lesheng Jin, Matei Zaharia, Joseph E. Gonzalez, and Sewon Min, plus CI, demo, and live status badges. A terminal-style box shows the pip install command, followed by code snippets for rendering a page to screenshots and searching a visual index of 8.28M Wikipedia pages.

    Ver todas las fotos de Tech

    ?

    Aún no hay comentarios. ¡Sé el primero!

    Más fotos de Tech

    Ver todas las fotos de Tech
    Elon Musk AGI warning vs nuclear weaponsElon Musk AGI warning vs nuclear weaponsIndian workers record jobs for robot trainingIndian workers record jobs for robot trainingVenus Aerospace Rotating Detonation Engine Test StandVenus Aerospace Rotating Detonation Engine Test StandiPhone 18 Pro and iPhone 18 Pro Max pricingiPhone 18 Pro and iPhone 18 Pro Max pricingApple Samsung crease memeApple Samsung crease memeSwitzerland switches to open-source2Switzerland switches to open-sourceApple iPhone gaming controllerApple iPhone gaming controllerNew Designers 2026New Designers 2026unique soccer team momentunique soccer team momentGoogle Workspace voice featuresGoogle Workspace voice featuresElon Musk AI warningElon Musk AI warningTesla next-generation RoadsterTesla next-generation RoadsterOriginal content program application reviewOriginal content program application reviewSamsung Galaxy Z Fold8 Tim Cook ad2Samsung Galaxy Z Fold8 Tim Cook adiPhone Duo pricing tiers2iPhone Duo pricing tiersNikola Tesla solitude quoteNikola Tesla solitude quote