Portrait of Adrian Bulat

Adrian Bulat

Adrian Bulat

Research ScientistSamsung AI Center Cambridge

Adrian Bulat is a Research Scientist at Samsung AI Cambridge. Previously, he received his PhD from the University of Nottingham where he worked with Dr. Georgios Tzimiropoulos as part of the Computer Vision Laboratory. His current research interests lie at the intersection of Computer Vision and Machine Learning.

Selected publications

Full publication list
  • 2026

    What Matters for Latent Reasoning with Flow Matching New

    Yassine Ouali, Adrian Bulat, Georgios Tzimiropoulos

    arXiv, 2026

    [abstract]

    FLaRe studies how language models can reason with continuous latent states and generate only their final answers. It combines a compact representation of symbolic reasoning, question-conditioned flow matching, and training on verified model-generated thoughts. The work evaluates whether these thoughts improve answers, support varied reasoning paths, admit faithful explanations, benefit from additional computation, and reduce inference cost. Experiments on arithmetic tasks examine the training choices needed to make latent reasoning effective.

    [cite]
    @article{ouali2026what,
      title={What Matters for Latent Reasoning with Flow Matching},
      author={Ouali, Yassine and Bulat, Adrian and Tzimiropoulos, Georgios},
      journal={arXiv preprint arXiv:2610.06666},
      year={2026},
      eprint={2610.06666},
      archivePrefix={arXiv},
      primaryClass={cs.LG},
      url={https://arxiv.org/abs/2610.06666}
    }
  • 2026

    UltraViT: Latency-Optimized On-device Vision Encoder for Large Vision-Language Models New

    Ioannis Maniadis Metaxas*, Adrian Bulat*, Alberto Baldrati*, Anestis Zaganidis, Yassine Ouali, Hyeonuk Kim, Georgios Tzimiropoulos

    European Conference on Computer Vision (ECCV), 2026

    [abstract]

    UltraViT addresses the cost of visual encoding when deploying large vision-language models on edge devices. Its pyramidal encoder combines different spatial mixing operations, with architectural choices guided by measured device latency rather than computational proxies alone. A two-stage pre-training procedure first transfers detailed spatial representations through dense distillation and then applies generative supervision from a frozen language model trained with mixed capacity. This procedure develops the semantic grounding needed for subsequent multimodal alignment and is reported to outperform contrastive and self-supervised alternatives. Experiments show that the combination of device-aware architecture design and generative training improves the efficiency and performance of vision encoding for large vision-language models, exceeding encoder-focused baselines while achieving approximately 1.7 times their on-device speed.

    [cite]
    @inproceedings{metaxas2026ultravit,
      title={{UltraViT}: Latency-Optimized On-device Vision Encoder for Large Vision-Language Models},
      author={Metaxas, Ioannis Maniadis and Bulat, Adrian and Baldrati, Alberto and Zaganidis, Anestis and Ouali, Yassine and Kim, Hyeonuk and Tzimiropoulos, Georgios},
      booktitle={European Conference on Computer Vision (ECCV)},
      year={2026},
      eprint={2607.23373},
      archivePrefix={arXiv},
      primaryClass={cs.CV},
      url={https://arxiv.org/abs/2607.23373}
    }
  • 2026

    Hierarchical Image Tokenization for Multi-Scale Image Super Resolution New

    Isma Hadji*, Enrique Sanchez*, Adrian Bulat, Brais Martinez, Georgios Tzimiropoulos

    International Conference on Machine Learning (ICML), 2026

    [abstract]

    This work introduces HIT, an approach to image super-resolution that adapts visual autoregressive modeling to produce several output resolutions in one forward pass. Conventional residual quantization does not ensure that intermediate token scales correspond to image scales, limiting earlier methods to a fixed output resolution. HIT instead builds a hierarchy of image tokens with overlap between scales, encouraging consistent representations at each resolution. Training also incorporates a direct preference optimization objective that favors the high-resolution target over its low-resolution input, using only the paired training images. Together, these changes support a 300-million-parameter model, compared with the billion-parameter VARSR baseline, without requiring external training data. The resulting system reports state-of-the-art super-resolution performance while reducing model size and providing outputs at multiple scales.

    [cite]
    @inproceedings{hadji2026hierarchical,
      title={Hierarchical Image Tokenization for Multi-Scale Image Super Resolution},
      author={Hadji, Isma and Sanchez, Enrique and Bulat, Adrian and Martinez, Brais and Tzimiropoulos, Georgios},
      booktitle={International Conference on Machine Learning (ICML)},
      year={2026},
      eprint={2605.14891},
      archivePrefix={arXiv},
      primaryClass={cs.CV},
      url={https://arxiv.org/abs/2605.14891}
    }
  • 2026

    VISion On Request: Enhanced VLLM efficiency with sparse, dynamically selected, vision-language interactions New

    Adrian Bulat*, Alberto Baldrati*, Ioannis Maniadis Metaxas*, Yassine Ouali, Georgios Tzimiropoulos

    IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2026

    [abstract]

    Existing approaches for improving the efficiency of Large Vision-Language Models (LVLMs) are largely based on the concept of visual token reduction. This approach, however, creates an information bottleneck that impairs performance, especially on challenging tasks that require fine-grained understanding and reasoning. In this work, we challenge this paradigm by introducing VISion On Request (VISOR), a method that reduces inference cost without discarding visual information. Instead of compressing the image, VISOR improves efficiency by sparsifying the interaction between image and text tokens. Specifically, the language model attends to the full set of high-resolution visual tokens through a small, strategically placed set of attention layers: general visual context is provided by efficient cross-attention between text-image, while a few well-placed and dynamically selected self-attention layers refine the visual representations themselves, enabling complex, high-resolution reasoning when needed. Based on this principle, we first train a single universal network on a range of computational budgets by varying the number of self-attention layers, and then introduce a lightweight policy mechanism that dynamically allocates visual computation based on per-sample complexity. Extensive experiments show that VISOR drastically reduces computational cost while matching or exceeding state-of-the-art results across a diverse suite of benchmarks, and excels in challenging tasks that require detailed visual understanding.

    [cite]
    @inproceedings{bulat2026visor,
      title={VISion On Request: Enhanced VLLM efficiency with sparse, dynamically selected, vision-language interactions},
      author={Bulat, Adrian and Baldrati, Alberto and Metaxas, Ioannis Maniadis and Ouali, Yassine and Tzimiropoulos, Georgios},
      booktitle={IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)},
      year={2026}
    }

Research

His current research interests lie at the intersection of Computer Vision and Machine Learning, with work conducted on topics such as efficient neural networks (via bit quantization, network binarization and compression) and human analysis (face alignment/recognition/super-resolution and human pose estimation).

Teaching

  • 2016–2018
    Teaching Assistant · University of Nottingham

    G52CPP: 2nd-year C++ Programming. G53SEC: 3rd-year Network Security.