16 accepted papers
Large Multimodal Models (LMMs) exhibit shortfalls when interpreting images and, by some measures, have poorer spatial cognition than young children or animals. Despite this, they attain high scores on many popular visual benchmarks, with headroom rapidly eroded by surging model progress. To address …
Large Language Models (LLMs) are increasingly deployed in real-world applications that demand complex reasoning. To track progress, robust benchmarks are required to evaluate their capabilities beyond superficial pattern recognition. However, current LLM reasoning benchmarks often face challenges su…
Large multimodal models (LMMs) have exhibited proficiencies across many visual tasks. Although numerous well-known benchmarks exist to evaluate model performance, they increasingly have insufficient headroom. As such, there is a pressing need for a new generation of benchmarks challenging enough for…
As the context limits of Large Language Models (LLMs) increase, the range of possible applications and downstream functions broadens. In many real-world tasks, decisions depend on details scattered across collections of often disparate documents containing mostly irrelevant information. Long-context…
Large multimodal models (LMMs) have proven flexible and generalisable across many tasks and fields. Although they have strong potential to aid scientific research, their capabilities in this domain are not well characterised. A key aspect of scientific research is the ability to understand and inter…
Workpiece placement with respect to an industrial robot plays an important role in robotic manufacturing due to its influence on the configuration-dependent properties of industrial robots. Suboptimal placements of the workpiece may increase the required joint torques and decrease the dexterity of t…
Telerobotic systems combined with miniaturised snake-like or elephant-trunk robotic arms can improve the ergonomics and accessibility in minimally invasive surgical tasks such as knee arthroscopy. Such systems, however, are usually designed in a specific and integral approach, making it expensive to…
Travel over sloped terrain is difficult as an incline changes the interaction between each wheel and the ground resulting in an unbalanced load distribution which can lead to loss of traction and instability. This paper presents a novel approach to generating wheel rotation for primary locomotion by…
Multi-legged robots are effective at traversing rough terrain. However, terrains that include collapsible footholds (i.e. regions that can collapse when stepped on) remain a significant challenge, especially since such situations can be extremely difficult to anticipate using only exteroceptive sens…
This paper proposes a novel configurable wheel that exhibits desired properties of varied radius wheels. Positional manipulation of the centre hub is proposed and tested to achieve these desired characteristics of `virtual' wheels in a physical system. The centre hub is manipulated via the use of pn…
While Product of Exponentials (POE) formula has been gaining maturity in modeling the kinematics of a serial-link robot, the Denavit-Hartenberg (D-H) notation is still the most widely used due to its intuitive and concise geometric interpretation of the robot. This paper has developed an analytical…
The Denavit-Hartenberg (D-H) model and the product of exponentials (POE) model have been two popular methods for modeling the kinematics of a serial-link robot. While these two models are equivalent in essence, no study has revealed how to convert from the POE model to the D-H model. The conversion
Continuum robots are increasingly used in minimally invasive surgeries. To date, the concentric tube mechanism and the cable-driven mechanism have been two prevalent mechanisms for constructing continuum robots. As these two mechanisms complement each other, it is worth exploring the possibility of
Knee arthroscopy is a very challenging surgical procedure that would strongly benefit from systems that can continuously map the inside of the knee, localize the arthroscope and surgical tools, and control instruments using visual information. A fundamental requirement of most of these systems is th
Natural language provides a convenient means of communicating information, and as such, is an ideal medium for enabling nonexpert users to teach robots novel tasks. However, in order to take advantage of natural language, a series of challenges must first be overcome. These challenges include the ne